Round 4 · MicroSheet · Season 1 finale · 3 October 2026

Season 1 goes to @schwarzkopfb.

The finale was the hardest task of the four: a spreadsheet in one HTML file, columns A to J and rows 1 to 50, formulas with operators, parentheses, cell references, SUM and AVERAGE over ranges, recalculation through chains, and errors that stay in the cell. Three entries, the first round with more than two, and everyone on the board played.

22
seats claimed
6
opened the task
3
submitted
3
passed the smoke test

Result

#EntrantToolRoundEntry
1@schwarzkopfbOpenCode, GPT-6 Sol2-0open the app
2@P1s0Claude1-1open the app
3@BingBangBoomChatGPT, GPT-5.6 Sol0-2open the app

Three survivors means a round robin: everyone against everyone, three comparisons, no pairing luck. @schwarzkopfb beat @P1s0, close, and @BingBangBoom, clear. @P1s0 beat @BingBangBoom, clear.

Season 1, final

#EntrantRatingRecordRounds won
1@schwarzkopfb12303-12 and 4
2@P1s011623-21 and 3
3@BingBangBoom9070-3

Two wins each, and the head to head between the top two ends 2-1. The rating puts @schwarzkopfb ahead because two of @P1s0's three wins came against the same opponent the leader also beat, while @schwarzkopfb's only loss was the Round 1 comparison that started the season.

How the finale was judged

Same probes as every round, run through the interface only, source never read: the five-item checklist with exact expected values (SUM over a range with an empty cell must read 190, a chain must update to 6 and 12, a circular reference and a garbage formula must show an error in their cell), then the task probes. Nested functions, a multi-column range, a range that contains a formula, operator precedence, a five-deep chain, circles through two and three cells, references outside the sheet, Escape and overtyping, copy and paste, undo, a reload, and fifty rows of formulas at once.

This round the three comparisons were run in a Claude Code session working the apps through the interface, instead of our usual judging script. To check that the change of judge did not change the outcome, the title-deciding pair was also run through the script used in Rounds 1 to 3. It picked the same winner, for the same reason, and also called it close.

One correction during the round. The first automated smoke run put @BingBangBoom out at 1 of 5. Three of the four failures said the tester ran out of actions before checking, and the fourth was because Tab does not move between cells in that app. The runbook says to check anything like that before eliminating it: a second automated run with a larger action budget passed 5 of 5, and so did a check by hand. The entry advanced.

The verdicts

A is the first entrant named, B the second.

@schwarzkopfb (B) over @P1s0 (A), close

Functional depth (40%) · B
Both engines returned identical, correct results on every probe: =SUM(A1:A3)+AVERAGE(B1:B3) 72, =SUM(A1:B3) 76, a range containing a formula cell 110, (A1+A2)*A3 1200 against A1+A2*A3 620, unary minus, 10/4 = 2.5, a five-deep chain to 1205, and 50 rows of formulas summing to 1375 and 2750 after A1 changed. A also evaluates MIN, MAX and COUNT. B has no functions beyond SUM and AVERAGE, but copy and paste rewrites references (C1 =A1+B1 pasted into C2 becomes =(A2+B2) and shows 30), Fill down does the same, and undo restores a changed value. A has none of copy, fill or undo, so every formula in a 50-row model has to be typed by hand.
UX & polish (25%) · A
A explains every error in plain words under the formula bar ('Divided by 0', 'No such cell: Z999 is outside the sheet (A1 to J50)', 'Unknown function: FOO is not available. Use SUM, AVERAGE, MIN, MAX or COUNT', 'A closing parenthesis is missing', 'Part of a loop through B1') and has a designed, intentional look. B shows bare codes, using #ERROR! for both an unknown function and a missing parenthesis, and a yellow 'Restore version / Back to current' bar stays on screen from first load. Against A: it opens as a filled-in trip plan with no way to a blank sheet, and columns I and J sit off-screen to the right at 1280 px.
Robustness (20%) · B
Neither froze: direct, two-cell and three-cell circles all showed an error and recovered once broken, the page answered within milliseconds throughout, both keep the sheet across a reload, and typing over a filled cell replaces it in both. B also lets a user recover from a mistake, with undo and a version history; in A an overwritten cell can only be retyped or the sample trip reloaded.
Ambition (15%) · A
A turns the engine into a trip planner: per-person and per-day costs, budget left, unpaid shares and category bars all recalculate live from the cells, with saved trips. One coherent extra that makes the sheet memorable. B's extras (fill down and right, undo and redo, formats, multiple documents, version history, export) each work but are spreadsheet breadth rather than one standout.
Overall · B, close
The engines are equal: every required probe and every extra probe gave identical, correct numbers in both, neither froze on any circle, and both survive a reload. That puts the decision on what each built around the engine. A is the more distinctive app: plain-language errors that name their cause, three extra functions, and a trip planner that uses the sheet for something real. B is the better spreadsheet: copy, paste and fill rewrite references correctly and undo works, which is what lets someone build a 50-row model without retyping every formula, and the task's own judging notes name exactly these as the extras worth most. B's loose ends are cosmetic; A's gaps are in editing itself. Close, and it would not take much to turn it.

The cross-check on this pair agreed on the winner and on why, and differed on two dimensions: it also gave polish and ambition to @schwarzkopfb, noting that its status line spells out a circular reference by name and that its grid fits the window. It credited @P1s0 with the most human-readable error message either app produced.

@P1s0 (A) over @BingBangBoom (B), clear

Overall · A, clear
Both engines are correct and both have a real idea on top. The separation is everything around the engine. A behaves like a spreadsheet when you use it: arrow keys, overtype that replaces, plain-language errors, three extra functions, and nothing lost on reload. B's editing has two faults a user hits within minutes, keyboard navigation that does not move and typing that appends to existing values, and one that silently changes data: a value committed in the formula bar is written again into the next cell clicked. B's X-ray and audit are the best explanation of formula dependencies in the field, and they deserve a working editor under them.

@schwarzkopfb (A) over @BingBangBoom (B), clear

Overall · A, clear
Equal engines again, so it comes down to whether the sheet is pleasant and safe to edit. A is: keyboard navigation, overtype, copy, fill and undo all behave as a spreadsheet user expects and nothing is lost on reload. B has the more original idea, making dependencies visible, and wins ambition for it, but its editing has a data-changing bug in the formula bar and two navigation faults that the X-ray cannot make up for.

How the winner got there

@schwarzkopfb's session ran seven prompts in OpenCode, an agent with separate plan and build modes, mostly on GPT-6 Sol. The first prompt went in about twenty minutes after the drop, a first version was submitted at 16:08 UTC, and the final one at 16:23. Prompts are quoted as typed.

prompt 1 · plan mode
let's build a small side project. first plan the arvhitecture and the implementation. de base spec is the following: [pasted the task]
Plan mode again, as in Round 2: the agent cannot write code yet, so the first output is an architecture.
prompt 2 · still planning
let's polish this plan: make it as user friendly and professional as possible. let's try to deliver a subset of what we have in google sheets but without the onl;ine services. also plan a local storage based multi-file storage system and per dociment revision control.
The prompt that decided the round. "A subset of Google Sheets" sets the bar at a real spreadsheet, not at the brief. The refined plan had a section titled "Editing should behave like a spreadsheet", and that is where copy, paste, fill and undo came from. The multiple documents and the version history came from the second half of the same sentence.
prompt 3 · build
let's build it based on the plan!
One instruction, and the first patch took eleven and a half minutes of agent time.
prompts 4 and 5
I've installed puppeteer on this system. now let's continue
ok, installed it locally in this repo. now you'd be able to import it
The agent wanted to test the app in a real browser and could not. Instead of skipping the tests, the human installed the tool it needed, twice, until it worked. Around thirty build and test steps followed.
prompt 6 · back to planning
could you add some helping UX features to make it more clear hoe to use this? eg. some hint on supported functions and the usage of them or the ability to select ranges with the cursor when editing the formula? like start typing SUM( and at that point the user selects a range in the grid and that selected range is added after the function like SUM(A1:A5 to help the user creating formulas. features like this
With a working, submitted sheet, the human switched back to plan mode for the next feature rather than asking for it straight away, and described it the way a user experiences it.
prompt 7
let's implenet it
The final version was resubmitted at 16:23 UTC, seven minutes before the close.

What the others did

@P1s0 spent the first part of the session questioning the brief: "it seems a simple excel workaround", "Should i build it first in excel, or is it just waste of time?", and then "Let come up with an environment where 500 cell serves more value then just a spreadsheet." That produced the trip planner. Then the same habits that won Rounds 1 and 3: a detailed build prompt, and a stream of requests written from the user's side, "when this happens i dont wanna see #VALUE!, oR formula error", which is where the plain-language errors came from. Submitted at 16:17. The trip planner was the most original idea of the round. What it gave up was the spreadsheet editing underneath, and in a spreadsheet task that cost it the title by a small margin.

@BingBangBoom did what worked last week: a first version submitted at 15:45, then forty minutes of testing as a user with the clock read out loud, "we have 27 mins left", "we have 17 mins left". The reasoning was sharp too. "I keep returning to this in the task rules: The formula engine is what separates entries. let's lean into that" produced the dependency X-ray and the formula audit. The testing found real gaps (no multi-cell selection, a save dialog with no way out, a formula helper that missed cells with formulas in them) but not the formula-bar bug that decided both of this entry's comparisons. Resubmitted at 16:27.

What we learned

All three engines were correct. We wrote the finale to separate entries on the formula engine, and it did not: every entry gave identical, correct numbers on every probe, and none froze on a circular reference. As in Round 3, the current tools handle the hard mechanics. What separated the field was what each person asked for around them, and whether the editing held up when used.

The ceiling you set is the ceiling you get. "A subset of Google Sheets" produced copy, fill, undo and version history. "An environment where 500 cells serve more value" produced a trip planner. "Lean into the engine" produced a dependency X-ray. Three good prompts, three different apps, and the task decided which one counted most.

Season 1 in four lines

Round 1 was won by an hour of iteration after the first build. Round 2 by fourteen minutes of planning, while the runner-up's build ran 38 minutes past the close. Round 3 by asking what the data was hiding. Round 4 by asking for a real spreadsheet rather than the brief. Two wins each for the top two, a newcomer in the last two rounds, and every result published with its receipts.

Twenty-two seats, three people who entered, and a format we now know works. Thank you to the three of you. Season 2 will be announced here.

Final leaderboard Round 3 write-up Round 1 write-up