Baseline study / 01 · 2026.10.10

Can basic networks
learn these
silhouettes?

Yes. A small CNN and ViT both learn to recognize Pokémon from noisy silhouettes. The CNN converges sooner; the ViT needs longer training. Larger changes in outline remain challenging.

1,025-WAY CLASSIFICATIONFROM SCRATCHMODERATE NOISESEED 17
Actual held-out noise sample
Clean reference silhouette of Pikachu

Clean reference

Actual moderate-noise test image of Pikachu

Moderate test input

CNN · familiar-outline test Top-1
1,025Target species / classes
6,383Training images · clean + noisy
96 × 96Network input · grayscale
0.098%Uniform-random Top-1 accuracy
01

What does the model actually see?

One image in, one species prediction out. These are real inputs and recorded predictions. Names, types, filenames and battle information are not supplied as model features.

01 / CLEAN SOURCEClean source silhouette for the noisy example

Original silhouette: 255 × 256.
This reference is in the training set.

02 / PUBLISHED NOISENoisy image from the published dataset

Blur, downsampling, contrast changes,
soft occlusion and random noise.

03 / NETWORK INPUTActual 96 by 96 grayscale network input

Bilinear resize to 96 × 96.
Normalize pixels to [−1, 1].

04 / PREDICTION

Ground truth · #0025

Recorded test predictions; this report does not run live inference.

Eight fixed species from the familiar-outline test set, selected independently of prediction success. The clean image explains the source; each model receives only one image at a time.

02

Two questions. Two evaluations.

Recognizing a familiar outline under new noise is different from recognizing another outline. We evaluate both, and select checkpoints using validation data only.

A / Familiar outline

Known reference, new noise

The clean reference outline is available during training. Test images use separate noise instances. This approximates recognition with a public reference gallery.

SHARED PARENT IS INTENTIONAL
B / Held-out outline

Withheld outline, clean or noisy

One outline hash per species is withheld, including all its clean and noisy descendants. None enter training or validation. The species itself is still represented in training.

EXACT OUTLINE-HASH OVERLAP = 0
SplitImagesSpeciesPurpose

Source data: 3,172 clean silhouettes plus the first 10 published moderate-noise shards (10,250 images). Training contains 2,142 clean and 4,241 noisy images. Some withheld outlines have no image in these ten noisy shards, so the held-out noisy test covers 980 species.

03

Small models, trained from scratch.

Both architectures start with random weights. We first compare 40-epoch runs, then add a 200-epoch ViT run after its training and validation curves reveal underfitting.

Small CNN

1.54M parameters
Conv 32→Conv 64→Conv 128→6 × 6 pool→256 → 1,025

Each block: 3 × 3 Conv → BatchNorm → ReLU → 2 × 2 MaxPool. The classification head uses 0.1 dropout.

Parameters
1,536,449
Training budget
40 epochs
Selected checkpoint
Epoch 39 · minimum validation cross entropy

Small ViT

0.95M parameters
8 × 8 patches→144 tokens→4 blocks→Mean pool→1,025

128-dimensional tokens, 4 attention heads, MLP width 512, learned positional embeddings and 0.1 dropout.

Parameters
952,321
Training budgets
40 epochs / additional 200-epoch run
Selected checkpoints
Epoch 40 / Epoch 171

Shared training setup

Optimizer
AdamW · LR 0.001 · weight decay 0.0001
Learning rate
3 warmup epochs, then cosine decay
Batch size
128
Loss
Cross entropy over 1,025 species
Augmentation
Scale 0.95–1.05, small translations, mild Gaussian grain
Stability
CUDA mixed precision · gradient clipping at 1.0
Environment
PyTorch 2.6.0+cu124 · RTX A5000

Training procedure

  1. Verify image hashes, labels and outline grouping.
  2. Check forward/backward passes and small-batch memorization.
  3. Train, evaluating familiar-outline validation data each epoch.
  4. Save the checkpoint with lowest validation loss.
  5. Evaluate test sets after training and save per-image predictions.

The longer ViT run starts again with the same random seed and a 200-epoch learning-rate schedule. It is not a continuation of the 40-epoch checkpoint, nor an equal-budget comparison.

04

When does learning happen?

The CNN learns familiar outlines quickly. The ViT is still underfit at 40 epochs, but reaches high accuracy with longer training. Curves come directly from saved epoch logs.

Familiar-outline validation · Top-1

CNN / 40 epViT / 40 epViT / 200 ep

Training accuracy includes augmentation and dropout; validation uses evaluation mode. These are not measurements under identical conditions.

Static learning curves and test results
05

Baseline results

The default metric is macro Top-1: compute accuracy for each species, then average equally across species. Species with more images therefore do not dominate the score.

Test accuracy · Macro Top-1

Model / training budgetFamiliar outline + noiseHeld-out outline + noiseHeld-out clean outline

Finding: The dataset contains learnable visual signal. Under this configuration, the CNN provides a strong baseline with fewer training epochs, while the longer-trained ViT also succeeds. This does not establish a universal architecture ranking or demonstrate learnability of battle outcomes.

06

What did the models predict?

These diagnostic cases are selected by outcome, not randomly sampled; they must not be used to estimate overall accuracy. Predictions compare the best CNN 40-epoch and ViT 200-epoch checkpoints.

07

How should we interpret the high scores?

A different outline hash does not necessarily mean a substantially different pose. We also compare each held-out clean image against its closest same-species training reference.

250 / 1,030

Held-out images closely resemble a reference

These images have binary silhouette IoU ≥ 0.95. Median nearest-reference IoU across the clean holdout is 0.898. No resized grayscale images match exactly, but near-duplicates remain.

Method: threshold at 128 on 96 × 96 inputs; take the maximum same-species clean-reference IoU, without extra alignment.

46.9% / 44.2%

Larger outline changes challenge both models

On 147 held-out clean images with IoU < 0.75, CNN / longer-trained ViT Top-1 falls to 46.9% / 44.2%. This post-training analysis did not alter the split or tune the models.

See actual reference / held-out image comparisonsSame-species training references and held-out outlines across similarity levels

Implication for the course challenge

With matching clean references, recognition on this moderate subset may be relatively easy. To assess pose generalization, near-duplicate outlines need further treatment. These results are an initial guide to difficulty.

What this pilot does not cover

One seed, one small noisy subset and two specific architectures. Hard / nightmare tiers, unseen species, pretrained models and battle prediction were not tested. Training has only 1–37 images per species.

Intrinsic ambiguity: Some Flapple / Appletun and Finizen / Palafin images share identical outlines. Their split assignments are coordinated to prevent cross-split overlap, but an image-only classifier still cannot always identify species uniquely.

08

Auditable and reproducible.

This report is generated from recorded experiment artifacts. All images, curves and result data are embedded. A downloaded copy opens offline in a browser.

Completed checks

  • Source SHA-256 checks and parent-label mapping
  • No held-out outline hashes in training or validation
  • All 1,025 classes represented in training
  • Finite losses and gradients for both architectures
  • 100% memorization on 32 images across seven species
  • Recorded Top-1 / Top-5 scores match saved predictions

Small-batch memorization checks the pipeline, not generalization. All descendants of an identical outline are grouped, but near-duplicates are not clustered into separate splits.

Experiment files

experiments/image_baselines/ ├── prepare.py # Data, label and split checks ├── train.py # CNN / ViT training and evaluation ├── audit_similarity.py # Outline similarity diagnostics ├── report.py # Static learning curves ├── build_html_report.py # Generate this report └── results/ # Metrics, logs, predictions, weights

Full commands are in the experiment README.md. Raw and prepared data are under datasets/. Checkpoints and per-image predictions are retained in the local run directories.

Reproduction commands and sources
# From the repository root, after setup and download (see README.md) .venv/bin/python experiments/image_baselines/prepare.py CUDA_VISIBLE_DEVICES=6 .venv/bin/python experiments/image_baselines/train.py --model cnn --out experiments/image_baselines/results/cnn_seed17 CUDA_VISIBLE_DEVICES=6 .venv/bin/python experiments/image_baselines/train.py --model vit --out experiments/image_baselines/results/vit_seed17 CUDA_VISIBLE_DEVICES=6 .venv/bin/python experiments/image_baselines/train.py --model vit --epochs 200 --out experiments/image_baselines/results/vit_seed17_200epochs .venv/bin/python experiments/image_baselines/build_html_report.py

Adjust GPU selection for your machine. Existing run directories are never overwritten: use new output paths and update report inputs for reruns. Short runs shared a GPU; elapsed times are not a fair speed benchmark.

Source repository: Shdakudow / pokemon-silhouette-dataset · Base revision 3d48401 · Moderate release: noise-moderate-v1. Original image credits are recorded in data/image_credits.csv and ACKNOWLEDGMENTS.md. Original artwork rights remain applicable.

Download metrics JSONDownload data audit JSONDownload learning curvesDownload offline HTML