We asked AI this question two different ways
“Brief us on portfolio performance: which practices are underperforming, what is driving it, and what we should do about it.”
Pasted verbatim into both. Separate sessions, same model, same settings.
TURNING DISTRESS INTO DISCIPLINE · CONTEXT DEMONSTRATION
Eighteen months of financials, time records and practice analytics from a twelve-practice dental group. Once submitted to an LLM (Claude) without context, the other received the same data plus seven markdown files describing the leadership team and its plan. Nothing else differed.

The gap is not model quality. It is that the model has your numbers and none of the judgment that makes them mean anything.
Everyone on your team can paste a spreadsheet into a model and get a confident, well-written answer back in seconds. That part is solved.
The model does not know how you define margin, what counts as bad, who owns the number, what already happened last quarter, or where a decision gets made. So it fills the gaps — fluently, and often wrongly.
Seven plain-language files, written once by the people who already carry this knowledge, sitting next to your data. The same model, the same question — a usable answer instead of a plausible one.
Not a dataset and not a tool — the judgment a leadership team already exercises, written down. Six layers that live in people's heads and in no system of record you own.
What does this number actually mean here?
Margin on collections, not production. Without it, a model — or a new hire — computes something defensible and wrong, and never signals doubt.
How bad is bad, and at what point does someone act?
Reappointment under 75% is red. A number with no threshold is an observation; a number with one is an alarm.
Who answers for this, by name?
Without a name, a finding is addressed to “management” — which means it is addressed to no one.
What are we actually trying to do this year?
The targets, and what is explicitly out of scope. This is what makes findings rankable and drift measurable.
What already happened that explains this?
A planned leave. Two resignations. Without the log, every variance gets investigated from scratch and half of them turn out to be nothing.
Where does this get decided, and when?
The third Tuesday. This is what turns a recommendation into a date on a calendar.
CASE STUDY · A TWELVE-PRACTICE DENTAL GROUP
Two uploads. The numbers in them are identical, byte for byte. The second adds roughly five thousand words that contain almost no data — definitions, thresholds, ownership, priorities, and what has already happened.
We did nothing clever. We opened a normal Claude chat, attached the files below straight into the message — the four spreadsheets on their own for the first run, then the same four spreadsheets plus the seven context files for the second — and typed the same question both times.
No tools, no plugins, no fine-tuning, no special instructions. Two fresh chats, the same model, the same prompt. The only thing that changed between them is what was attached to the message.
4 files. About 1,600 rows. No context of any kind. Click any file to open it.
The same 4 files, plus 7 markdown files — roughly 5,000 words and not one additional number. Click any file to open it.
Both runs were asked through Claude. Nothing about the result depends on which model you use — the same two uploads and the same prompt run the same way on any general-purpose model.
Cedar Ridge Dental Partners is a fictional Dental Group. Lets break down some interesting finds when we look at the data submitted to a LLM before adding a context layer.
On paper, the second-best profit margin in the group this quarter.
MisleadingProduction fell by half, into an operating loss for the quarter.
MisleadingThe highest collection rate in the group, at 102%.
MisleadingBought in January. Thin margin, a third of its bills unpaid past ninety days.
MisleadingOpened in February. Losing money every month.
MisleadingAll five of those sentences are misleading.
That is the whole demonstration.
Every one of them is arithmetically correct. Not one of them leads a reader to the right conclusion. What separates them is implementing a context structure: seven files, written down by the people who run the business.
The question — practice performance
“Brief us on portfolio performance: which practices are underperforming, what is driving it, and what we should do about it.”
Pasted verbatim into both. Separate sessions, same model, same settings.
Five practices, same rows in both runs. Without context, the obvious reading of the numbers is misleading in every one.
Riverbend has three hygiene chairs. Two of its three hygienists resigned in February and March, and neither seat was refilled. Wages dropped immediately, so the margin climbed to 23.3% — second-best in the group. In the same six months the share of hygiene patients leaving with their next visit booked fell from 84% to 61%. That lost revenue shows up in roughly three quarters. A saw the margin and listed Riverbend with the healthy practices. B saw the recall collapse and opened the briefing with it.
“in the high-teens”
“and it is also our single biggest hidden risk”
“reappointment rate: 84.2% → 61.2%”
Stonebridge's production fell by about half. Its only full-time dentist was on parental leave from April to late June — notified in February, approved, with a locum covering three days a week and the staff deliberately kept on so the hygiene schedule would survive. A saw a collapse and called it the most urgent item in the briefing, recommending a second locum. B recognised a funded absence going exactly to plan, and recommended nothing at all.
“staff the locum (or a second locum) at full coverage”
“This is a funded absence performing as designed, not a performance problem.”
Oak Hollow posts the best collection rate in the group — 102%. That figure is collections divided by net production, meaning the total after write-offs. Oak Hollow writes off 37 cents of every dollar it bills, under an insurance contract signed below market rates, and that shrinks the denominator until the ratio looks superb. Measured against everything it actually billed, it collects 65 cents on the dollar — second-worst in the group. A printed the 102% in its scorecard and never mentioned the practice again.
Does not raise it
“but this is the write-off artifact our metric definitions specifically warn about”
“Gross collection rate is 64.7%”
Pinecrest was bought in January and is still being absorbed. Its margin is thin and a third of its unpaid bills are more than ninety days old. Both runs found the practice. A blamed front-desk process and proposed an audit. B knew Pinecrest is still billing on its old fee schedule — about 12% below the group's — which was meant to be migrated within 150 days of the purchase and never was. That one unmigrated file explains most of the gap, and it costs $12,000 to fix.
“the legacy front-office/billing processes and patient-engagement practices from before the acquisition appear not to have been converted”
“Legacy fee schedule (~12% below Cedar Ridge's) still not migrated”
“the single largest driver”
Harborview opened its doors in February and is losing money, which is what a brand-new practice does for its first year. A worked that out on its own, from the fact that Harborview's rows only begin partway down the file — a good inference. B additionally knew the practice is running ahead of the ramp plan the board approved, and that its heavy marketing spend is that plan rather than an overrun.
“this is a de novo practice in its fifth month of operation, ramping up on a normal trajectory”
“Collections tracking ahead of ramp plan”
“are the approved launch budget, not a variance”
Every line in a context brain is something a leadership team should already be able to say out loud. Writing it down is the management work; the model just makes its absence visible in about four seconds.
None of these are AI problems. They are what an auditor, a lender or a new CFO would find — and each is fixed by writing one page.
Seven files, about five thousand words, containing almost no data — one afternoon of a leadership team's time. Open any of them.
Write your own seven files.
The console walks your team through every question, shows the markdown as you go, and exports the whole folder when you are done.
Open the interview console