Pokémon / Project notes
October 10, 2026
Download HTML
Intro to ML · Challenge design

Pokémon ML Challenge

Train/test options, battle rule ideas, and initial CNN and ViT image-recognition results.

1. Train/test options

Proposed task: predict the winner from two noisy Pokémon images. The battle rule is still to be defined.

For either option, provide a complete reference gallery: clean images, species IDs, types and base stats for all Pokémon, including test species. Training pairs include outcomes; test outcomes are withheld.

Option 1

Split by Pokémon

Give students a complete Pokémon information list covering all training and test species: species IDs, clean reference images, types and base stats.

Training and test battles use different species.

TRAINSpecies A, B, C, DExample pairs: A–B, A–C, B–D, C–D
TESTSpecies E, F, G, HExample pairs: E–F, E–G, F–H

Tests: transferring battle rules to Pokémon with no training battle outcomes.

Students can use the full information list to identify and look up test Pokémon. Their information is public; their battle outcomes are withheld. This tests generalization to species with no training battle records.

Option 2

Split by pair

Training and test battles can use the same species, but different matchups.

TRAINA–B, A–C, B–D, C–DAll four species have training battle records.
TESTA–D, B–CNeither matchup appears in training.

Tests: predicting new matchups between familiar Pokémon.

Usually an easier starting point. Give each species enough varied training opponents so the model can learn useful interactions.

A–H represent different species. Keep A–B and B–A, and all image/noise versions of that matchup, in the same split. Validation should follow the chosen split strategy. Both options are proposals; neither has been evaluated for battle prediction yet.

2. Battle rule ideas

Three candidate rules for generating battle labels. These are simplified teaching rules, not the official Pokémon battle system; none has been evaluated in the image experiments below.

A. Type-aware score · simplest

Calculate a weighted strength from each Pokémon’s base stats, then adjust it by its type effectiveness against the opponent. Compare the two adjusted scores; the higher score wins.

Learning task: combine strength and type advantage. Useful as a first baseline, but check that raw stat totals alone do not solve most matchups.

B. Short deterministic duel · suggested starting point

Start with HP from the public stats. Damage depends on physical attack versus defense, or special attack versus special defense, plus type effectiveness. Use one fixed attack-selection policy; the faster Pokémon attacks first. Repeat until a knockout or a fixed round limit.

Learning task: learn interactions between offense, defense, speed and type. No random misses or critical hits. Equal speed means simultaneous attacks; at the round limit, compare remaining HP fractions. Exact ties are draws and can be excluded if we choose a binary challenge. Damage scaling, attack policy and round limit still need calibration.

C. Duel with a public arena · optional extension

Add an arena label, such as neutral, sunny or rainy, that changes selected type bonuses. Students receive two images + the arena label; the same matchup may have a different winner in a different arena.

Learning task: predict context-dependent matchups. For a pair split, keep every arena, image variant and reversed ordering of the same pair in one partition.

Suggested plan: start with B and use A as a simple comparison. First test learnability using true Pokémon stats under both split strategies; then add image recognition. Compare against majority-class and total-stat baselines, and check that opponent interactions matter.

Decisions before generating labels
  • Make outcomes deterministic and independent of left/right ordering. Swapping the pair must swap the winner; define one consistent draw policy.
  • Provide the complete Pokémon information list in both options. If the exact generator is also public, students can identify the Pokémon and simulate the result. To emphasize learning battle relationships, keep the exact generator on the organizer side and provide labelled examples.
  • Calibrate rules using training/validation data, then freeze the generator before producing the final test set. Check class balance and data volume; these ideas still need a separate battle-prediction pilot.

3. CNN & ViT results

Image recognition only: these experiments predict a species from one image. They do not evaluate either battle-prediction split above.

Swipe the table horizontally to see all models.

Test set / metricCNN
40 epochs
ViT
40 epochs
ViT
200 epochs

Macro Top-1 gives equal weight to each species; Top-1 gives equal weight to each image. The first row is balanced, so both metrics agree. Uniform-random Top-1 is about 0.098%.

What does “larger silhouette differences” mean? For each held-out clean image, compare it with every clean training reference of the same species. If even the closest reference has silhouette overlap (IoU) below 0.75, include it in the last row. There are 147 such images. This measures outline differences, which may reflect pose, form, proportions or position—not a manually labelled pose change.

4. Image input

One grayscale image goes in; the model predicts the Pokémon species. No names, types, filenames or battle information are used as input features.

Clean reference

Actual clean Pikachu reference silhouette

255 × 256 · available in training for this example.

New noisy test image

Actual moderate-noise test image of Pikachu

Published moderate noise: blur, reduced contrast, soft occlusion and grain.

Actual network input

Actual 96 by 96 grayscale model input

Resized to 96 × 96; pixel values normalized to [−1, 1].

A fixed illustrative test example. Predictions are recorded outputs, not live inference.

5. Training

6,383 training images: 2,142 clean references + 4,241 noisy images. Both models use random initialization, AdamW (initial LR 0.001), batch size 128, mild image augmentation and a cosine learning-rate schedule.

CNN: 3 convolution blocks, 1.54M parameters. ViT: 4 transformer blocks with 8 × 8 patches, 0.95M parameters. The 40-epoch ViT was underfit; a separate 200-epoch run shows that it can learn with a longer budget.

CNN · 40 epochsViT · 40 epochsViT · 200 epochs

Training loss

Familiar-outline validation accuracy

Checkpoints are selected by minimum validation loss, never test accuracy. The 200-epoch ViT uses a longer learning-rate schedule and starts from scratch; it is not an equal-budget comparison.

6. Takeaways

This is a one-seed pilot on a small moderate-noise subset. It does not test unseen species, hard/nightmare noise, or prove that CNNs are always better than ViTs.

Split details and limitations
  • Source: 3,172 clean images plus 10,250 published moderate-noise images. All 1,025 species appear in training.
  • Familiar-outline validation/test: 1,025 images each, one per species. Their clean references are available during training.
  • Held-out outline: one exact outline hash per species is withheld, including all clean/noisy descendants. This gives 1,030 clean images (1,025 species) and 3,959 noisy images (980 species).
  • No withheld outline hash appears in training or validation. Near-duplicates remain: 250 of the 1,030 clean holdouts have reference IoU ≥ 0.95.
  • Similarity is measured on 96 × 96 binary silhouettes, threshold 128, without extra alignment. The 147-image analysis is post-training and was not used for tuning.
  • Some different species share identical silhouettes; image-only identity can therefore be ambiguous.