Product case study · Jacob Dixon

NoSchnitz

A five-handed Sheepshead game for a group of friends who could never all be free on the same Thursday — and a working demonstration of how I run product: research before sequence, measurement before belief, and every reason written down where the next person can check it.

0numbered releases, each with written rationale
0days from first commit to the release this is measured at
0issues filed on a labelled Now/Next/Later board
0named test and measurement harnesses

Live at noschnitz.com. Solo against four AI opponents, or multiplayer with a shareable link. Figures below are as of v0.120.1.

01 — Where the roadmap came from

It started with an interview, not an idea

The group had already tried this. Understanding why it failed was the whole roadmap.

During the pandemic the same five-to-seven people ran a biweekly Thursday Sheepshead night on an existing site, with a separate video call alongside it, coordinated over group text. It failed more often than it worked.

The interview found that it wasn't failing for lack of interest. It was failing because a scheduled commitment kept losing to spontaneous life, and getting five people to commit in advance was the actual bottleneck. On the good nights the opposite problem appeared: six or seven people wanted in, and the tool couldn't gracefully rotate anyone through five seats.

The finding that reordered everything was smaller and less flattering to the product: the video call was doing the emotional work. The card game was an excuse to fill the quiet spaces.

Which meant the thing to build first was not a better card game. It was the removal of the coordination tax — and a social layer treated as first-class rather than a second tab.

That became a written hypothesis at the top of the roadmap, with everything in Now scoped to test it and nothing else: will this group actually play together, on a whim, without a scheduled event? AI-filled seats so there is no minimum headcount. A shareable link so nobody has to cajole a group text. Presence so "is it worth hopping on" answers itself.

Why this matters more than the artifact it produced

A roadmap that starts from a hypothesis can be wrong in public. Six days in, the first real session ran two humans and three AI seats through roughly twenty hands with no stalls — the table filled to five with only two people in the room, which is precisely the bet. The write-up in the repo says so and then immediately says not to read it as coverage: two humans is not five, nobody joined mid-hand, nobody backgrounded a phone. A result you are willing to under-claim is a result you can still learn from.

02 — Sequence

The roadmap is an argument, not a list

A board shows what is being worked on and never why that rather than something else.

Features decompose into epics and epics into user stories, with stable IDs so a commit can name the story it implements. But the part that does the work is the paragraph under each bucket explaining why the order is the order — because a reader who can't see the reasoning concludes the sequencing was accidental rather than chosen.

Now

Scoped to test one hypothesis: will they play on a whim, unscheduled?

Epic · MP-1Shareable table linksOne click to a table, a link you can text, live arrivals.
Epic · MP-2AI seat-fillPlayable the instant one human shows up. No minimum headcount.
Epic · MP-3Guest joinA name and a seat. No account, no download.
Epic · COM-3Seamless seat coverStep away, hand the seat to AI, reclaim it. One AFK player can never freeze a table.
Bug · area: aiClairvoyance leak in the endgame solverThe solver carried the real partner into every sampled world. Filed with the reproduction, not the symptom.

Next

Earns its place only once the Now hypothesis has an answer.

Epic · MP-5Table lifecycleLeaving, watching, coming back — including from a second device.
Epic · AI-1Variant-aware judgementThe engine's bars were tuned under one rule set and are applied under several.
Epic · AI-2Pick qualityThe marginal hand loses money. Named as an epic so the fix isn't a constant nudge.
StoryReal-device multiplayer coverageEvery multiplayer bug worth fixing came from a person on a phone, not a harness.
area: toolingA daily report that carries informationIts version note fires nearly every day, so it says nothing. A signal that never varies is not a signal.

Later

Deliberately deferred — with the reason attached, so deferral stays a decision.

Epic · COM-1Embedded voice/videoLaunch-critical by the research, paused by a decision recorded on the epic itself.
Epic · COM-2PresenceDepends on persistent accounts to be more than in-app.
StoryCollapse the table POSTsNot cleanup — the API sits at eleven of the platform's twelve function slots. This is what makes room to build anything.
Protected scopeLearn-to-play tutorialNot low priority. It only pays off once there is somewhere to send a new player.

Protected scope, and what it costs to protect

The tutorial is core to the long-term goal and sits in Later anyway, because it can't pay off until public tables exist. Deferring it would normally mean quietly killing it — so the deferral came with an architecture constraint written into the same document: multiplayer had to be built as an additional mode alongside solo, reusing the pure rules engine, rather than a rewrite that entangles local and networked state. The tutorial slots in later exactly the way solo works today.

Sequencing a thing later is easy. Keeping it cheap to build later is the actual work.

03 — A decision, in full

When the measurement lost

The number was unambiguous. It was also answering a different question than the one that mattered.

In Sheepshead the picker can go alone — no partner, quadruple stakes. The engine holds a threshold for when the AI takes that option. The app holds a second threshold for when it offers the option to a human. There is no rule that says they must be the same number.

Across 20,239 hands where the picker had a partner available, going alone at a threshold of 17 was behind by 1.9 points per hand, negative in four of four seeds. It only turned positive at 18. The offer bar was set to 18 on the strength of exactly that.

Then somebody played the game.

Evidence
20,239 pickers with a partner available. Alone trails by 1.9 points per hand at 17; negative in 4 of 4 seeds; positive at 18.
Optimises
Expected units per hand. The correct objective, measured correctly.
What a player sees
An opponent goes alone on a hand their own screen refused to offer them. The game appears to know something it will not tell them.
Failure mode
A trust cost that no harness in the repo can measure, traded for a gain that every harness can.

Shipped: 17.v0.46.0
Overriding a measurement is only defensible when you say what it costs, write the condition that would reverse you, and make the codebase fail loudly if someone later moves half of it.

"What should the AI do?" and "what should we offer a human?" are different questions. The answer is a product call sitting on top of a measurement — and reading the two as one number is how a team optimises itself into a game nobody enjoys.

04 — Evidence

Roughly forty instruments, and three ways they lie

Every one of these rules exists because a specific measurement was wrong first.

Engine and AI changes are decided by paired A/B runs: a variant in one seat against four unchanged seats over identical shuffles, across several seeds. What makes a small result trustworthy is the null control — an arm identical to the baseline must return exactly +0.0000, and the test suite asserts it. A harness that can't return zero when nothing changed can't be believed when something does.

That was the easy part. The hard-won part is the standing list of ways the instruments mislead:

Dilution A rule that fires rarely reads as nothing in a whole-hand aggregate. Two separate rules measured flat — then +0.252 and +0.210 per firing once the denominator was the decisions the seat actually faced. One of them cost a full measurement pass concluding "not established" from a diluted number that was, undiluted, a 4.5-sigma effect. Check the firing rate before concluding anything
The wrong harness A one-seat A/B structurally cannot see cooperation: one defender stands down, the other two still contest the trick, and somebody else pays for the saving. When two harnesses genuinely disagreed, the resolution was not to prefer one — it was to build the third that answered the actual question. Don't settle a disagreement by picking a favourite
Fixtures that search A test that deals hands until it finds the shape it wants is sampling a fresh population every run. One did that at a measured 17% failure rate from the day it was written, went red on the main branch for a commit that had passed on its own PR, and skipped a production verification. Nothing was wrong with the release. Run a suite a dozen times before believing it is stable

The rule underneath all three

Aggregate simulation is the safety net, not the detector. Every AI fix that has actually landed started from one hand a human flagged, was reproduced against the engine before anything changed, and was then pinned as an assertion with a negative control. Several correct-looking diagnoses measured as pure noise.

05 — Cadence

212 releases in 29 days

Real data from the changelog. Each bar is one day; each unit is one versioned release with a written rationale.

Jul 24Aug 7Aug 21
Releases that day Day a decision below was made
Interview with the group. The roadmap is written from it the same day.
First real multiplayer session — two humans, three AI seats, ~20 hands, no stalls. The seat-fill bet pays; the write-up refuses to call it coverage.
Multiplayer promoted to production, verified by fetching the live bundle and grepping it for the version string — never by build hash, which the platform's minifier makes non-deterministic.
First corpus mining run over 131 recorded hands: no large systematic gap against a competent human. Subtler signal needs an order of magnitude more hands, and that becomes a filed issue rather than a conclusion.
Named, durable tables ship — along with a written list of the three things deliberately not built, each with what would change the answer.
The backlog file is retired one day after it was created; the board moves to issues. The two-agent boundary is written down.
Cadence change: every merge no longer ships to production. An integration branch goes in between, and promotion becomes a deliberate act.
Chat ships — the first community feature — and reaches production a promotion earlier than intended through a flag scope accident. Caught by the automated content check, and blessed rather than rolled back, on the record.

06 — The operating system

What the process actually is

Two AI agents work this repo under a written boundary — one owns the roadmap and the board, one owns the code, and I own the decisions. Everything below exists because that only works if the contract is explicit.

  1. One fact, one home. Cite it; never restate it.The test is simple: if this number changes, how many files must I edit? More than one is a defect, not a convenience. The project's own handoff doc refuses to restate the version number — because it once claimed v0.7.2 while the app shipped v0.22.0, and a stale number reads exactly like a fresh one.
  2. Every reference must be openable by whoever has to act on it.An issue that cites a rule by number, in a document the implementer cannot open, is not specifying. It is appealing to authority. Cite repo-relative paths, or inline the part being relied on and stop citing the rest.
  3. Say what is explicitly out of scope.And name the drift, not the principle. "Keep it focused" prevents nothing; "do not also move this constant — that is a separate measured change" prevents the thing that actually happens. This is the single highest-yield habit on the board.
  4. Pre-register the measurement.Instrument, null control, sample size, splits, engine version — and a control arm that would catch your own approximation. Deciding how you'll know you're right before doing the work is the whole difference between a measurement and a rationalisation.
  5. Separate the tell from the objective.A symptom that is easy to measure will get optimised if you let it. State in the issue which number is the symptom and which is the goal, where someone under time pressure will read it.
  6. Mark what is not shaped yet.An unshaped note is fine. An unshaped note wearing a shaped note's clothes is not — the reader can't tell "not thought through" from "small", and guesses wrong.
  7. Corrections stay visible.Status changes require evidence; a wrong conclusion gets a follow-up marked superseded, not a silent edit. Same for the role-based lenses the project reasons through: each carries a dated learnings log, and an untested learning is filed as a hypothesis rather than a fact.

The shared file that lasted one day

The backlog started as a markdown file. Within twenty-four hours it had collected a duplicate heading, two shipped items still open, and a line describing a deleted file as present in the working tree. It was retired for the issue board — not on taste, but on the observation that every mechanism a shared mutable file needs — section ownership, a claim protocol, re-read-before-closing, batching to control cost — is scaffolding for a merge conflict. Issues have no merge.

07 — The uncomfortable one

Rigour drifts to where rigour is cheap

The most useful document in this project is the one that audits the rest of it.

Partway through, the board got read the way an outsider would read it. The finding was structural rather than careless, and it is the one I'd want a hiring committee to see.

The AI pick threshold — an item worth a fraction of a point per hand — had a null control, sign agreement across splits, a throw-in budget check and a control arm under a second rule set. Real-device multiplayer coverage sat in Later with no acceptance criteria at all and a note that it was "largely covered in practice." By the project's own written evidence, every multiplayer bug worth fixing had come from a person on a phone, and the newest, least-exercised surfaces had never been touched by five people at once.

The measurable thing got measured. That is not the same as the important thing getting measured.

The output wasn't a reshuffle. It was a standing question, written into the conventions where it gets asked again: is this rigorous because it matters most, or because it was the easiest thing to be rigorous about?

The related test for whether reasoning is genuinely written down: hand the board to someone who wasn't in the conversation. Everything they need is reachable from a clone. Every number has one home. Every gate can be evaluated by someone who wasn't there. Every issue says what not to do. Anything unshaped is labelled unshaped. Where that fails, the reasoning is still in somebody's head — with a paper trail that looks like documentation, which is the more expensive failure, because it's the one nobody checks.

08 — Constraints

The platform's limits are roadmap inputs, not chores

The API sits at eleven of the hosting plan's twelve serverless function slots, and the limit only fails at deploy time. That single fact is why "collapse the action endpoints" is a filed, prioritised backlog item framed as what makes room to build anything rather than as cleanup — and why any change that wants a new route reads that constraint first.

Deployments, not CI minutes, are the scarce resource: the plan allows a hundred a day, and one day of two-agent work spent well over thirty. So the deploy cost of every branch is written down as a table — zero per push to a working branch, one per merge to the integration branch, one per promotion — and documentation-only changes skip the build via a version-controlled ignore rule whose every failure direction is toward building rather than toward a missed deploy.

And the release process changed when the cost model did

For most of the project every merge shipped to production. Once real people were playing between releases, that stopped being right: an integration branch went in the middle, and promotion became a deliberate, milestone-driven act with an explicit merge strategy — because squashing a promotion loses the merge base and makes the next promotion re-conflict everything already shipped. The old cadence isn't deleted from the record; it's marked historical, so a future reader knows which instructions are current.

09 — What this is evidence for

Two claims