AI read the lease but could not divide a pie
ChatGPT read a 45-page lease in three minutes for Dr. Nikki Siso, then failed at recipe arithmetic. The failures sort by information, not difficulty.
The same model that read a 45-page commercial lease in three minutes could not divide a pie into twelve slice costs. Dr. Nikki Siso, founder of Conscious Kitchen in Westlake Hills outside Austin, used ChatGPT successfully for lease review, oven sourcing and herbal formulation, then failed on spreadsheet arithmetic and on a permitting sequence that cost her $30,000. The failures do not sort by difficulty.
They sort by something else, and the line is the useful part of this interview for anyone building.
What actually predicted success here?
Every task ChatGPT did well for Siso was a research task with a publicly documented answer. Lease terms are written down. Oven specifications are published. Herbal formulations appear in the literature. The mechanism of food preservatives is documented science. In each case the model did strong work quickly, in a domain where the operator had no training.
Every task it failed was a process or craft task where the answer is local, tacit or structural. The permitting sequence lives in contractors' heads. Reliable spreadsheet manipulation is a structural weakness. Design taste is what a professional charges for. That is Artificial Reality's read of the pattern, not Siso's phrasing — but it is drawn entirely from the tasks she walks through.
The lease worked. The pie did not.
Prompted to act as a commercial real estate agent, ChatGPT identified eight terms in a 45-page lease that Siso should negotiate and drafted the letter to her landlord. An attorney reviewed the draft in ten minutes for free instead of charging $1,000.
The costing task was arithmetic: a pound of almonds at $24, one cup per recipe, twelve slices per pie, updated when almond prices move.
"I had one hell of a time getting it to actually fill in the spaces on the Excel chart. For some reason it did not want to fill in the blanks. Or it would just duplicate and it would be the wrong values."
"That was not well orchestrated. For some reason it was a major challenge and it did not do what I wanted it to do. It would have taken me many more hours to perfect it."
Galina Fendikevich put it to her directly that AI had created more work than doing it by hand. Siso's answer: "Yes."
Was this a prompting problem?
No, and that matters more than the failure itself. Siso, who has no technical background, independently arrived at role assignment as her method — act like an electrician, act like a commercial real estate agent, act like a herbologist — for each of the three tasks she describes in detail. The technique is well known among people who follow this closely. That she reached it unaided, and that it was sufficient, suggests the gap between expert and novice users may be narrower than the discourse implies.
She was not held back by not knowing how to prompt. She was held back by information that does not exist in the machine, and by arithmetic reliability the model did not deliver.
What this changes for a roadmap
Three things follow, and they are uncomfortable in useful ways.
- Demo difficulty is a bad proxy for value. Reading a forty-five-page lease is objectively harder than knowing to commission a mechanical engineering plan before applying for a permit. The model did the hard one brilliantly and missed the easy one entirely, because the easy one was never written down anywhere.
- The moat is undocumented process, not model quality. Where the answer lives in a contractor's head or on an inspector's laminated form, no general-purpose model reaches it. Whoever collects that information owns a product a frontier lab does not automatically eat.
- Spreadsheet reliability is a product, not a feature request. The failure was not conceptual. It was duplicated rows and wrong values on unit conversion and division by yield, in a task the operator could describe precisely and check.
The commercially interesting part
Siso's successes cost her nothing and saved her $1,000 of legal billing. Her failures cost real money — $30,000 on the permitting sequence — and neither had a product she could buy at her price. She evaluated purpose-built recipe costing software at around $150 a month and found it more than her budget allowed.
So the demand is documented at both ends and unserved in the middle: a competent operator who already prompts well, already has the general tool, and still has two expensive problems it will not solve.
What to do with this
Run the test on your own category. List the last ten tasks your product was used for and sort them into research with a public answer versus process with a local answer. If everything sits in the first column, a general-purpose model is your competitor and price is your only defence. If you can move into the second column — by collecting the sequence, the form, the local rule that nobody has written down — that is where an operator who lost $30,000 will pay you.
Which raises the question worth answering before your next roadmap review: what does your best customer know that no model can look up?
Where this comes from
S1E6: I trusted AI to build my restaurant, it cost me $30,000 — the full interview with Dr. Nikki Siso. Listen or watch: YouTube, Spotify or Apple Podcasts.