STICK ARENA / RELEASE 05 / 6 September 2026

A better spectacle.
An honest test of evolution.

Watch, skip and replay now use the same seeded combat. Weapons have distinct illustrated silhouettes. Fighters negotiate terrain, families remain recognisable, and the ecology is tested hundreds of generations beyond its founders.

← Return to the arena
116,660PRODUCTION BRACKET & INVITATIONAL GAMES
THROUGH GENERATION 300
8,800RELEASE WEAPON BALANCE BOUTS
BOTH STARTING SIDES
1,056WEAPON / STYLE / ELEMENT COVERAGE BOUTS
CONTACT, PACING & MECHANICS

Does experience survive inheritance?

Three timelines reached generation 300: two exploratory populations of 64 and an untouched validation population of 256. Final populations beat their founders in 78.9%, 81.8% and 79.9% of bouts. The default population won 57.8% against its generation 50 and 50.0% against generation 100. Progress is substantial but can plateau; these finite runs do not establish perpetual improvement.

Across every generation from 200 through 300, effective weapon counts averaged 4.00, 4.17 and 7.86; the largest weapon averaged 45.2%, 41.7% and 25.1% of its population. Seven, eight and ten different weapons won championships in those late windows. Several simpler diversity policies failed strength or variety checks and were rejected.

Independent default-population checks returned 82.0% against fresh broad profiles, 55.5% against its generation 50 and 53.5% against generation 100. A 47.7% result against generation 200 had an approximate paired interval of 43.2–52.1%; that comparison establishes neither improvement nor the fixed five-point noninferiority margin. Initial estimates are retained, and no parameter was tuned to the validation seed.

An independent 1,024-pair comparison of both exploratory populations against the original selection policy’s generation-300 cohort returned 48.1% [45.9–50.4%] and 52.2% [49.9–54.5%]. Both satisfy the predeclared five-percentage-point noninferiority margin. These comparisons do not establish absolute superiority over that particular opponent population.

Abilities and elements still develop common combinations. Natural generation-200 and generation-300 fights used all five abilities across the audit and in both default-size cohorts, with an effective 2.44–3.26 skill types and blast/shield most common. Across 1,152 bouts there were no time-limit decisions and one quiet gap above eight seconds; its recorded poses show an active ice exchange. The game preserves selection and counters, without promising permanent equilibrium or universal excitement.

Measured checkpoints; lines connect samples. On a narrow screen, scroll horizontally (left/right keys when focused), or open the full chart. Every plotted value is also available below.

Win rates at sampled checkpoints against frozen founding and generation-50 populations. Shaded bands show approximate paired 95 percent intervals. Read every checkpoint in the table below.
Read every measured checkpoint
Timeline and generationvs foundersvs gen 50vs gen 100
Exploratory AGeneration 1072.1%Approx. 95%: 67.9%–76.2%256 pairs · 512 bouts
Exploratory AGeneration 2572.1%Approx. 95%: 68.3%–75.9%256 pairs · 512 bouts
Exploratory AGeneration 5069.1%Approx. 95%: 64.7%–73.6%256 pairs · 512 bouts
Exploratory AGeneration 10079.7%Approx. 95%: 75.9%–83.4%256 pairs · 512 bouts59.4%Approx. 95%: 55.0%–63.8%256 pairs · 512 bouts
Exploratory AGeneration 20073.2%Approx. 95%: 69.2%–77.3%256 pairs · 512 bouts57.2%Approx. 95%: 52.9%–61.6%256 pairs · 512 bouts49.0%Approx. 95%: 44.4%–53.6%256 pairs · 512 bouts
Exploratory AGeneration 30078.9%Approx. 95%: 75.1%–82.7%256 pairs · 512 bouts54.3%Approx. 95%: 49.7%–58.9%256 pairs · 512 bouts54.5%Approx. 95%: 49.8%–59.1%256 pairs · 512 bouts
Exploratory BGeneration 1070.7%Approx. 95%: 66.6%–74.8%256 pairs · 512 bouts
Exploratory BGeneration 2568.8%Approx. 95%: 64.5%–73.0%256 pairs · 512 bouts
Exploratory BGeneration 5065.4%Approx. 95%: 60.9%–69.9%256 pairs · 512 bouts
Exploratory BGeneration 10073.6%Approx. 95%: 69.8%–77.5%256 pairs · 512 bouts57.8%Approx. 95%: 53.3%–62.3%256 pairs · 512 bouts
Exploratory BGeneration 20081.4%Approx. 95%: 78.0%–84.9%256 pairs · 512 bouts67.2%Approx. 95%: 62.7%–71.6%256 pairs · 512 bouts54.5%Approx. 95%: 49.8%–59.2%256 pairs · 512 bouts
Exploratory BGeneration 30081.8%Approx. 95%: 78.3%–85.4%256 pairs · 512 bouts69.5%Approx. 95%: 65.3%–73.8%256 pairs · 512 bouts62.3%Approx. 95%: 58.0%–66.7%256 pairs · 512 bouts
Default 256 validationGeneration 1069.7%Approx. 95%: 65.9%–73.5%256 pairs · 512 bouts
Default 256 validationGeneration 2573.6%Approx. 95%: 69.6%–77.6%256 pairs · 512 bouts
Default 256 validationGeneration 5077.5%Approx. 95%: 73.9%–81.2%256 pairs · 512 bouts
Default 256 validationGeneration 10082.8%Approx. 95%: 79.3%–86.3%256 pairs · 512 bouts52.5%Approx. 95%: 48.0%–57.0%256 pairs · 512 bouts
Default 256 validationGeneration 20080.7%Approx. 95%: 77.2%–84.1%256 pairs · 512 bouts53.7%Approx. 95%: 49.0%–58.4%256 pairs · 512 bouts50.6%Approx. 95%: 45.9%–55.3%256 pairs · 512 bouts
Default 256 validationGeneration 30079.9%Approx. 95%: 76.4%–83.4%256 pairs · 512 bouts57.8%Approx. 95%: 53.4%–62.2%256 pairs · 512 bouts50.0%Approx. 95%: 45.7%–54.3%256 pairs · 512 bouts
Generation 300vs foundersvs gen 50vs gen 100Late weapon diversity
Exploratory A64 fighters · seed 9704378.9%Approx. 95%: 75.1%–82.7%256 pairs · 512 bouts54.3%Approx. 95%: 49.7%–58.9%256 pairs · 512 bouts54.5%Approx. 95%: 49.8%–59.1%256 pairs · 512 bouts4.0045% largest weapon share
Exploratory B64 fighters · seed 18490781.8%Approx. 95%: 78.3%–85.4%256 pairs · 512 bouts69.5%Approx. 95%: 65.3%–73.8%256 pairs · 512 bouts62.3%Approx. 95%: 58.0%–66.7%256 pairs · 512 bouts4.1742% largest weapon share
Default 256 validation256 fighters · seed 42957179.9%Approx. 95%: 76.4%–83.4%256 pairs · 512 bouts57.8%Approx. 95%: 53.4%–62.2%256 pairs · 512 bouts50.0%Approx. 95%: 45.7%–54.3%256 pairs · 512 bouts7.8625% largest weapon share

Intervals use side-swapped matchup pairs as the units of replication. They describe matchup-sampling uncertainty conditional on each saved population, not variation between independent evolutionary runs. Comparisons with the original-policy baseline concern those saved cohorts; they do not establish general algorithm-level superiority. These runs use actual Tournament methods: champion placement, best-of-three finals, automatic historical invitationals and their heirs, selection, crossover, mutation and immigrants. Rendering, ceremony time and manual timeline wars are omitted. 456,456 additional historical selection bouts are counted separately. Evaluation outcomes and evaluation arena seeds never feed selection. Champion training and evaluation populations can share ancestry. The founder reference is a population sample, not a guarantee against every earlier fighter. Developed generations can cycle or plateau; capped stat budgets do not promise endless power growth.

Effective weapon count is exp(Shannon entropy); a perfectly even four-weapon pool scores 4, while a near-monoculture approaches 1. Diversity columns average generations 200–300. Appearance and weapon labels alone do not establish interesting combat.

Read the matchups, not just the average.

Each cell shows a row weapon’s wins against the column weapon. Paired profiles share their non-weapon genes, isolating weapon choice; separate neutral profiles set those genes to 0.5. Finite samples expose strengths and counters, without proving that every optimized build or future ecosystem is balanced.

Side-swapped games are correlated. A harsh matchup can be a useful counter if the disadvantaged weapon has viable answers elsewhere. Evolved-profile stress tests, specialist counterplay and actual population outcomes accompany these broad matrices; a 50% random-build average alone is insufficient.

Space for the spectacular.

MechanicEvents
Wall runs694
Wall kicks171
Wall bounces124
Launches1,462
Air hits567
Spikes372
Weapon recalls132
Incoming-weapon catches21
Cover breaks255

The coverage sweep includes all eleven weapon choices, six inherited attack preferences, four elements and rotating signature skills. This is a pacing and occurrence test; its unmatched opponents do not form a balance matrix.

0 decisions in 1,056 bouts. The 95th-percentile duration was 44.2 seconds. The 95th-percentile longest interval without a hit, block, parry or bind was 5.9 seconds. Dodges, near misses and movement can still be active during that interval.

Recalls require a returning weapon; catches require incoming thrown steel and a poised unarmed recipient. Their rarity creates punctuation, rather than a mandatory ritual in every bout. Direct fixtures also verify ownership changes, reset resources, solid cover and shield overflow. Team and three-fighter exhibitions receive separate completion checks.

What the evolved population actually does

Separate natural pairings among late-generation fighters test the strategies that selection actually produced. These draws do not force a weapon mix. Action kinds count distinct observed move categories per bout, including defensive actions; they are evidence of variety rather than a direct measure of enjoyment.

Living cohortMean duration95th-percentile quiet gapLongest quiet gapAction kinds / boutDecisions
64 fighters · seed 97043Generation 200 · 192 bouts28.5s5.3s7.8s16.00
64 fighters · seed 97043Generation 300 · 192 bouts28.7s4.9s5.7s16.70
64 fighters · seed 184907Generation 200 · 192 bouts29.7s5.3s7.3s15.70
64 fighters · seed 184907Generation 300 · 192 bouts29.3s5.7s9.9s15.00
256 fighters · seed 429571Generation 200 · 192 bouts30.7s5.1s7.9s17.40
256 fighters · seed 429571Generation 300 · 192 bouts28.0s5.0s6.6s16.70

A family worth following.

Hair, clothes, faces, palettes, crests and small personal details pass through an independent visual genome. Children resemble their parents while siblings have their own details. Illustrated weapons retain their materials, strings, elemental inlays and both daggers in full replays.

Open Fighters to follow a family, meet its descendants and rewatch a recent bout. Open Gene Pool → Ancestral Trials to test your own timeline against a saved earlier reference. Trials display conservative 95% bounds over independent matchup pairs and never train your population. The interactive trial uses a bounded-score Hoeffding interval, which stays nonzero even after a perfect sample; the larger offline curves above use the separately labeled approximate paired intervals.

The living population, family identities, champions, current bracket and followed families survive saves. Completed bracket history is retained for the current session. Old fights are explicitly restaged under the new combat version. When browser storage fails, a visible notice offers export; export serializes the current in-memory timeline.

Reviewed from both sides.

Adversarial reviews challenged combat equivalence, mechanics, evolution claims, optimized balance, weapon art, persistence, navigation, accessibility and spectator flows. Separate refuters attempted to disprove the findings and fixes. Confirmed defects were repaired; unsupported recommendations, including a blanket slam nerf, were rejected after broader tests.

Release checks cover watched/headless equivalence at multiple frame rates, skipped and resumed finals, actual browser FX and KO replays, corruption-safe import, storage failure, ancestry and family persistence, mobile and landscape layouts, keyboard dialogs and reduced-motion rendering. Watched and background paths agree within a runtime. Tiny floating-point differences between JavaScript runtimes can grow into different seeded outcomes, so shared recipes reproduce fighters and arenas rather than promising cross-browser identical results. Long native runs measure distributions under the shared rules; they are not transcripts of a particular browser timeline. Empirical readiness is a body of reproducible evidence, not a claim that software is incapable of failure.