The finale was the hardest task of the four: a spreadsheet in one HTML file, columns A to J and rows 1 to 50, formulas with operators, parentheses, cell references, SUM and AVERAGE over ranges, recalculation through chains, and errors that stay in the cell. Three entries, the first round with more than two, and everyone on the board played.
| # | Entrant | Tool | Round | Entry |
|---|---|---|---|---|
| 1 | @schwarzkopfb | OpenCode, GPT-6 Sol | 2-0 | open the app |
| 2 | @P1s0 | Claude | 1-1 | open the app |
| 3 | @BingBangBoom | ChatGPT, GPT-5.6 Sol | 0-2 | open the app |
Three survivors means a round robin: everyone against everyone, three comparisons, no pairing luck. @schwarzkopfb beat @P1s0, close, and @BingBangBoom, clear. @P1s0 beat @BingBangBoom, clear.
| # | Entrant | Rating | Record | Rounds won |
|---|---|---|---|---|
| 1 | @schwarzkopfb | 1230 | 3-1 | 2 and 4 |
| 2 | @P1s0 | 1162 | 3-2 | 1 and 3 |
| 3 | @BingBangBoom | 907 | 0-3 |
Two wins each, and the head to head between the top two ends 2-1. The rating puts @schwarzkopfb ahead because two of @P1s0's three wins came against the same opponent the leader also beat, while @schwarzkopfb's only loss was the Round 1 comparison that started the season.
Same probes as every round, run through the interface only, source never read: the five-item checklist with exact expected values (SUM over a range with an empty cell must read 190, a chain must update to 6 and 12, a circular reference and a garbage formula must show an error in their cell), then the task probes. Nested functions, a multi-column range, a range that contains a formula, operator precedence, a five-deep chain, circles through two and three cells, references outside the sheet, Escape and overtyping, copy and paste, undo, a reload, and fifty rows of formulas at once.
This round the three comparisons were run in a Claude Code session working the apps through the interface, instead of our usual judging script. To check that the change of judge did not change the outcome, the title-deciding pair was also run through the script used in Rounds 1 to 3. It picked the same winner, for the same reason, and also called it close.
A is the first entrant named, B the second.
The cross-check on this pair agreed on the winner and on why, and differed on two dimensions: it also gave polish and ambition to @schwarzkopfb, noting that its status line spells out a circular reference by name and that its grid fits the window. It credited @P1s0 with the most human-readable error message either app produced.
@schwarzkopfb's session ran seven prompts in OpenCode, an agent with separate plan and build modes, mostly on GPT-6 Sol. The first prompt went in about twenty minutes after the drop, a first version was submitted at 16:08 UTC, and the final one at 16:23. Prompts are quoted as typed.
let's build a small side project. first plan the arvhitecture and the implementation. de base spec is the following: [pasted the task]
let's polish this plan: make it as user friendly and professional as possible. let's try to deliver a subset of what we have in google sheets but without the onl;ine services. also plan a local storage based multi-file storage system and per dociment revision control.
let's build it based on the plan!
I've installed puppeteer on this system. now let's continue
ok, installed it locally in this repo. now you'd be able to import it
could you add some helping UX features to make it more clear hoe to use this? eg. some hint on supported functions and the usage of them or the ability to select ranges with the cursor when editing the formula? like start typing SUM( and at that point the user selects a range in the grid and that selected range is added after the function like SUM(A1:A5 to help the user creating formulas. features like this
let's implenet it
@P1s0 spent the first part of the session questioning the brief: "it seems a simple excel workaround", "Should i build it first in excel, or is it just waste of time?", and then "Let come up with an environment where 500 cell serves more value then just a spreadsheet." That produced the trip planner. Then the same habits that won Rounds 1 and 3: a detailed build prompt, and a stream of requests written from the user's side, "when this happens i dont wanna see #VALUE!, oR formula error", which is where the plain-language errors came from. Submitted at 16:17. The trip planner was the most original idea of the round. What it gave up was the spreadsheet editing underneath, and in a spreadsheet task that cost it the title by a small margin.
@BingBangBoom did what worked last week: a first version submitted at 15:45, then forty minutes of testing as a user with the clock read out loud, "we have 27 mins left", "we have 17 mins left". The reasoning was sharp too. "I keep returning to this in the task rules: The formula engine is what separates entries. let's lean into that" produced the dependency X-ray and the formula audit. The testing found real gaps (no multi-cell selection, a save dialog with no way out, a formula helper that missed cells with formulas in them) but not the formula-bar bug that decided both of this entry's comparisons. Resubmitted at 16:27.
All three engines were correct. We wrote the finale to separate entries on the formula engine, and it did not: every entry gave identical, correct numbers on every probe, and none froze on a circular reference. As in Round 3, the current tools handle the hard mechanics. What separated the field was what each person asked for around them, and whether the editing held up when used.
The ceiling you set is the ceiling you get. "A subset of Google Sheets" produced copy, fill, undo and version history. "An environment where 500 cells serve more value" produced a trip planner. "Lean into the engine" produced a dependency X-ray. Three good prompts, three different apps, and the task decided which one counted most.
Round 1 was won by an hour of iteration after the first build. Round 2 by fourteen minutes of planning, while the runner-up's build ran 38 minutes past the close. Round 3 by asking what the data was hiding. Round 4 by asking for a real spreadsheet rather than the brief. Two wins each for the top two, a newcomer in the last two rounds, and every result published with its receipts.
Twenty-two seats, three people who entered, and a format we now know works. Thank you to the three of you. Season 2 will be announced here.