AI Products
Shipping an AI feature in six weeks without regretting it
The order I build in, the guardrails I set on day one, and the questions I ask before picking a model.
August 2, 2026 · 2 min read
Six weeks is enough time to ship an AI feature that works, if you spend the time in the right order. Most teams spend it in the wrong order. They start with the interface, then the model, then discover in week five that the answers are not good enough and there is no way to measure why. Here is the order I use instead.
Week one is all evals
Before any interface exists I write twenty to fifty examples of what good looks like, using the inputs the real product will see. Real customer emails, real product descriptions, real documents, with the output I would be happy to ship next to each one.
That set becomes the test suite for every prompt and model change afterwards. It is the single highest leverage hour of the project, and it is the one most often skipped because it does not feel like building.
Week two: the thinnest possible loop
Input in, model call, output on screen. No accounts, no styling, no history. Just the core loop, wired to the fastest model to set up. Run the eval set through it. Look at the failures. Most of what you learn about the product you learn here, in a plain text box, for almost no money.
Weeks three and four: design for the wrong answer
Now the interface, and the first thing it gets is a path for when the model is wrong. An edit control, a retry, a way to flag the answer, a visible source where there is one. Streaming helps too. People forgive a slow answer they can watch forming far more than a fast one they have to wait for blindly.
Accounts, limits and billing come in here as well, because a model call has a cost per use, and a product without limits is a product with an unbounded bill.
Week five: pick the model
Only now. Run the eval set against two or three candidates and compare quality, speed and cost. Often a smaller, cheaper model wins on the actual task. Because the product was built so the provider can be swapped, this is an afternoon, not a rewrite.
Week six: harden and launch
Logging, monitoring, cost alerts, a kill switch, moderation where user content is involved. Then launch to a small group, watch the logs for a week, and widen. A month of support follows, because the first real users will find the cases the eval set missed, and those cases go into the set.
The feature is not the model. The feature is what happens around it.
Why this order works
Every step produces something you can measure before you spend on the next one. If the thin loop in week two cannot produce good answers, you know in week two, not week five, and the money is still in the bank. That is the whole trick. Not better models, just a better order.
Next article
A website should explain the product in ten seconds
Why I start every site with one sentence and three proofs, and how that shapes everything below the fold.
Read