STICK ARENA / RELEASE 05 / 6 September 2026A better spectacle.
An honest test of evolution.
Watch, skip and replay now use the same seeded combat. Weapons have distinct illustrated silhouettes. Fighters negotiate terrain, families remain recognisable, and the ecology is tested hundreds of generations beyond its founders.
← Return to the arena
116,660PRODUCTION BRACKET & INVITATIONAL GAMES
THROUGH GENERATION 300
8,800RELEASE WEAPON BALANCE BOUTS
BOTH STARTING SIDES
1,056WEAPON / STYLE / ELEMENT COVERAGE BOUTS
CONTACT, PACING & MECHANICS
Does experience survive inheritance?
Three timelines reached generation 300: two exploratory populations of 64 and an untouched validation population of 256. Final populations beat their founders in 78.9%, 81.8% and 79.9% of bouts. The default population won 57.8% against its generation 50 and 50.0% against generation 100. Progress is substantial but can plateau; these finite runs do not establish perpetual improvement.
Across every generation from 200 through 300, effective weapon counts averaged 4.00, 4.17 and 7.86; the largest weapon averaged 45.2%, 41.7% and 25.1% of its population. Seven, eight and ten different weapons won championships in those late windows. Several simpler diversity policies failed strength or variety checks and were rejected.
Independent default-population checks returned 82.0% against fresh broad profiles, 55.5% against its generation 50 and 53.5% against generation 100. A 47.7% result against generation 200 had an approximate paired interval of 43.2–52.1%; that comparison establishes neither improvement nor the fixed five-point noninferiority margin. Initial estimates are retained, and no parameter was tuned to the validation seed.
An independent 1,024-pair comparison of both exploratory populations against the original selection policy’s generation-300 cohort returned 48.1% [45.9–50.4%] and 52.2% [49.9–54.5%]. Both satisfy the predeclared five-percentage-point noninferiority margin. These comparisons do not establish absolute superiority over that particular opponent population.
Abilities and elements still develop common combinations. Natural generation-200 and generation-300 fights used all five abilities across the audit and in both default-size cohorts, with an effective 2.44–3.26 skill types and blast/shield most common. Across 1,152 bouts there were no time-limit decisions and one quiet gap above eight seconds; its recorded poses show an active ice exchange. The game preserves selection and counters, without promising permanent equilibrium or universal excitement.
Measured checkpoints; lines connect samples. On a narrow screen, scroll horizontally (left/right keys when focused), or open the full chart. Every plotted value is also available below.

Read every measured checkpoint
| Timeline and generation | vs founders | vs gen 50 | vs gen 100 |
|---|
| Exploratory AGeneration 10 | 72.1%Approx. 95%: 67.9%–76.2%256 pairs · 512 bouts | — | — |
|---|
| Exploratory AGeneration 25 | 72.1%Approx. 95%: 68.3%–75.9%256 pairs · 512 bouts | — | — |
|---|
| Exploratory AGeneration 50 | 69.1%Approx. 95%: 64.7%–73.6%256 pairs · 512 bouts | — | — |
|---|
| Exploratory AGeneration 100 | 79.7%Approx. 95%: 75.9%–83.4%256 pairs · 512 bouts | 59.4%Approx. 95%: 55.0%–63.8%256 pairs · 512 bouts | — |
|---|
| Exploratory AGeneration 200 | 73.2%Approx. 95%: 69.2%–77.3%256 pairs · 512 bouts | 57.2%Approx. 95%: 52.9%–61.6%256 pairs · 512 bouts | 49.0%Approx. 95%: 44.4%–53.6%256 pairs · 512 bouts |
|---|
| Exploratory AGeneration 300 | 78.9%Approx. 95%: 75.1%–82.7%256 pairs · 512 bouts | 54.3%Approx. 95%: 49.7%–58.9%256 pairs · 512 bouts | 54.5%Approx. 95%: 49.8%–59.1%256 pairs · 512 bouts |
|---|
| Exploratory BGeneration 10 | 70.7%Approx. 95%: 66.6%–74.8%256 pairs · 512 bouts | — | — |
|---|
| Exploratory BGeneration 25 | 68.8%Approx. 95%: 64.5%–73.0%256 pairs · 512 bouts | — | — |
|---|
| Exploratory BGeneration 50 | 65.4%Approx. 95%: 60.9%–69.9%256 pairs · 512 bouts | — | — |
|---|
| Exploratory BGeneration 100 | 73.6%Approx. 95%: 69.8%–77.5%256 pairs · 512 bouts | 57.8%Approx. 95%: 53.3%–62.3%256 pairs · 512 bouts | — |
|---|
| Exploratory BGeneration 200 | 81.4%Approx. 95%: 78.0%–84.9%256 pairs · 512 bouts | 67.2%Approx. 95%: 62.7%–71.6%256 pairs · 512 bouts | 54.5%Approx. 95%: 49.8%–59.2%256 pairs · 512 bouts |
|---|
| Exploratory BGeneration 300 | 81.8%Approx. 95%: 78.3%–85.4%256 pairs · 512 bouts | 69.5%Approx. 95%: 65.3%–73.8%256 pairs · 512 bouts | 62.3%Approx. 95%: 58.0%–66.7%256 pairs · 512 bouts |
|---|
| Default 256 validationGeneration 10 | 69.7%Approx. 95%: 65.9%–73.5%256 pairs · 512 bouts | — | — |
|---|
| Default 256 validationGeneration 25 | 73.6%Approx. 95%: 69.6%–77.6%256 pairs · 512 bouts | — | — |
|---|
| Default 256 validationGeneration 50 | 77.5%Approx. 95%: 73.9%–81.2%256 pairs · 512 bouts | — | — |
|---|
| Default 256 validationGeneration 100 | 82.8%Approx. 95%: 79.3%–86.3%256 pairs · 512 bouts | 52.5%Approx. 95%: 48.0%–57.0%256 pairs · 512 bouts | — |
|---|
| Default 256 validationGeneration 200 | 80.7%Approx. 95%: 77.2%–84.1%256 pairs · 512 bouts | 53.7%Approx. 95%: 49.0%–58.4%256 pairs · 512 bouts | 50.6%Approx. 95%: 45.9%–55.3%256 pairs · 512 bouts |
|---|
| Default 256 validationGeneration 300 | 79.9%Approx. 95%: 76.4%–83.4%256 pairs · 512 bouts | 57.8%Approx. 95%: 53.4%–62.2%256 pairs · 512 bouts | 50.0%Approx. 95%: 45.7%–54.3%256 pairs · 512 bouts |
|---|
| Generation 300 | vs founders | vs gen 50 | vs gen 100 | Late weapon diversity |
|---|
| Exploratory A64 fighters · seed 97043 | 78.9%Approx. 95%: 75.1%–82.7%256 pairs · 512 bouts | 54.3%Approx. 95%: 49.7%–58.9%256 pairs · 512 bouts | 54.5%Approx. 95%: 49.8%–59.1%256 pairs · 512 bouts | 4.0045% largest weapon share |
|---|
| Exploratory B64 fighters · seed 184907 | 81.8%Approx. 95%: 78.3%–85.4%256 pairs · 512 bouts | 69.5%Approx. 95%: 65.3%–73.8%256 pairs · 512 bouts | 62.3%Approx. 95%: 58.0%–66.7%256 pairs · 512 bouts | 4.1742% largest weapon share |
|---|
| Default 256 validation256 fighters · seed 429571 | 79.9%Approx. 95%: 76.4%–83.4%256 pairs · 512 bouts | 57.8%Approx. 95%: 53.4%–62.2%256 pairs · 512 bouts | 50.0%Approx. 95%: 45.7%–54.3%256 pairs · 512 bouts | 7.8625% largest weapon share |
|---|
Intervals use side-swapped matchup pairs as the units of replication. They describe matchup-sampling uncertainty conditional on each saved population, not variation between independent evolutionary runs. Comparisons with the original-policy baseline concern those saved cohorts; they do not establish general algorithm-level superiority. These runs use actual Tournament methods: champion placement, best-of-three finals, automatic historical invitationals and their heirs, selection, crossover, mutation and immigrants. Rendering, ceremony time and manual timeline wars are omitted. 456,456 additional historical selection bouts are counted separately. Evaluation outcomes and evaluation arena seeds never feed selection. Champion training and evaluation populations can share ancestry. The founder reference is a population sample, not a guarantee against every earlier fighter. Developed generations can cycle or plateau; capped stat budgets do not promise endless power growth.
Effective weapon count is exp(Shannon entropy); a perfectly even four-weapon pool scores 4, while a near-monoculture approaches 1. Diversity columns average generations 200–300. Appearance and weapon labels alone do not establish interesting combat.
Read the matchups, not just the average.
Each cell shows a row weapon’s wins against the column weapon. Paired profiles share their non-weapon genes, isolating weapon choice; separate neutral profiles set those genes to 0.5. Finite samples expose strengths and counters, without proving that every optimized build or future ecosystem is balanced.
Side-swapped games are correlated. A harsh matchup can be a useful counter if the disadvantaged weapon has viable answers elsewhere. Evolved-profile stress tests, specialist counterplay and actual population outcomes accompany these broad matrices; a 50% random-build average alone is insufficient.
Space for the spectacular.
| Mechanic | Events |
|---|
| Wall runs | 694 |
|---|
| Wall kicks | 171 |
|---|
| Wall bounces | 124 |
|---|
| Launches | 1,462 |
|---|
| Air hits | 567 |
|---|
| Spikes | 372 |
|---|
| Weapon recalls | 132 |
|---|
| Incoming-weapon catches | 21 |
|---|
| Cover breaks | 255 |
|---|
The coverage sweep includes all eleven weapon choices, six inherited attack preferences, four elements and rotating signature skills. This is a pacing and occurrence test; its unmatched opponents do not form a balance matrix.
0 decisions in 1,056 bouts. The 95th-percentile duration was 44.2 seconds. The 95th-percentile longest interval without a hit, block, parry or bind was 5.9 seconds. Dodges, near misses and movement can still be active during that interval.
Recalls require a returning weapon; catches require incoming thrown steel and a poised unarmed recipient. Their rarity creates punctuation, rather than a mandatory ritual in every bout. Direct fixtures also verify ownership changes, reset resources, solid cover and shield overflow. Team and three-fighter exhibitions receive separate completion checks.
What the evolved population actually does
Separate natural pairings among late-generation fighters test the strategies that selection actually produced. These draws do not force a weapon mix. Action kinds count distinct observed move categories per bout, including defensive actions; they are evidence of variety rather than a direct measure of enjoyment.
| Living cohort | Mean duration | 95th-percentile quiet gap | Longest quiet gap | Action kinds / bout | Decisions |
|---|
| 64 fighters · seed 97043Generation 200 · 192 bouts | 28.5s | 5.3s | 7.8s | 16.0 | 0 |
|---|
| 64 fighters · seed 97043Generation 300 · 192 bouts | 28.7s | 4.9s | 5.7s | 16.7 | 0 |
|---|
| 64 fighters · seed 184907Generation 200 · 192 bouts | 29.7s | 5.3s | 7.3s | 15.7 | 0 |
|---|
| 64 fighters · seed 184907Generation 300 · 192 bouts | 29.3s | 5.7s | 9.9s | 15.0 | 0 |
|---|
| 256 fighters · seed 429571Generation 200 · 192 bouts | 30.7s | 5.1s | 7.9s | 17.4 | 0 |
|---|
| 256 fighters · seed 429571Generation 300 · 192 bouts | 28.0s | 5.0s | 6.6s | 16.7 | 0 |
|---|
A family worth following.
Hair, clothes, faces, palettes, crests and small personal details pass through an independent visual genome. Children resemble their parents while siblings have their own details. Illustrated weapons retain their materials, strings, elemental inlays and both daggers in full replays.
Open Fighters to follow a family, meet its descendants and rewatch a recent bout. Open Gene Pool → Ancestral Trials to test your own timeline against a saved earlier reference. Trials display conservative 95% bounds over independent matchup pairs and never train your population. The interactive trial uses a bounded-score Hoeffding interval, which stays nonzero even after a perfect sample; the larger offline curves above use the separately labeled approximate paired intervals.
The living population, family identities, champions, current bracket and followed families survive saves. Completed bracket history is retained for the current session. Old fights are explicitly restaged under the new combat version. When browser storage fails, a visible notice offers export; export serializes the current in-memory timeline.
Reviewed from both sides.
Adversarial reviews challenged combat equivalence, mechanics, evolution claims, optimized balance, weapon art, persistence, navigation, accessibility and spectator flows. Separate refuters attempted to disprove the findings and fixes. Confirmed defects were repaired; unsupported recommendations, including a blanket slam nerf, were rejected after broader tests.
Release checks cover watched/headless equivalence at multiple frame rates, skipped and resumed finals, actual browser FX and KO replays, corruption-safe import, storage failure, ancestry and family persistence, mobile and landscape layouts, keyboard dialogs and reduced-motion rendering. Watched and background paths agree within a runtime. Tiny floating-point differences between JavaScript runtimes can grow into different seeded outcomes, so shared recipes reproduce fighters and arenas rather than promising cross-browser identical results. Long native runs measure distributions under the shared rules; they are not transcripts of a particular browser timeline. Empirical readiness is a body of reproducible evidence, not a claim that software is incapable of failure.