Thinking Like An Insurance Risk Manager
Issue 2 /
From the CEO's desk
In the first edition, we talked about Coverage Analysis as one piece of the end-to-end workflow a complex claims adjuster needs. The response surprised us, real feedback, real suggestions, enough of it that a second edition was an easy call.
We spend most of our time in our company on two questions: how casualty claims workflows actually run, and how risk managers actually think.
This edition is about the second one, specifically, what happens when AI tries to replicate a risk manager's brain by throwing more AI tokens at the problem.
The brain doesn't work that way, and neither should AI.
The human brain runs on roughly 20% of the body's energy while being about 2% of its weight. Children, still building out their concept, context, and knowledge graphs, spend 50 to 60%. Efficiency isn't a nice-to-have here, it's the whole design. Rote memory, concept learning, and contextual reasoning work together so that a decision costs only what it needs to.
Token-maxxing is the AI equivalent of energy waste. It shows up when a system doesn't know the context or the concept, so it substitutes brute-force search for judgment, and the bill lands on your compute budget.
Think about learning to bake your mother's cake recipe. You probably won't nail it on the first try. By the fifth, you're close, because each attempt builds on context and concept, not just repetition. No model gets there in five trials, or in anything close to it. Even after tens of millions of training trials, a model is still assembling its knowledge graphs and concept graphs from something closer to scratch. Extrapolate that gap to enterprise scale and the economics get worse, not better: build your own frontier model in-house, and the cost will outrun your annual revenue several times over before it outruns the problem.
Andrew Ng made a related point in a recent newsletter: past a certain threshold, increasing token usage gives diminishing returns because there are organizational bottlenecks that burning more tokens alone cannot resolve. He was equally direct about the incentive behind the hype, companies that sell tokens profit when you use more of them, the same instinct that has repair shops recommending 3,000-mile oil changes and toothpaste ads showing a six-inch ribbon when a pea-sized dab does the job.
None of this means tokens are bad. It means the win isn't in how many you use, it's in whether the system reasons the way a good risk manager would before it starts spending them.
Approaches We've Taken To Keep A Check On Token-maxxingWe changed the metric. Tokens used per month tells you nothing. We track tokens per feature per month, and correlate that against the value it generated for customers, not billings. Billings-as-proxy is how you fall into the token-maxxing trap: it looks fine as long as the customer keeps paying, until a more efficient competitor dislodges you. That's also not a responsible way to build. We instrument cost per unit of output. Once an application scales past a prototype, we track what it actually costs to run. One of our applications costs roughly 1 cent per page, but the variance is large, a page with 10 medical entities costs very differently from one with 100. That's a 10x swing on information density alone. Knowing this number lets us do fast back-of-envelope math before we build, not after. We preserve optionality. We maintain a model garden and test across providers continuously. Even at the prototype stage, we architect for the possibility of switching providers, including open-weight options, or build the first version to run against multiple LLMs at once. This is advice no single model provider can credibly give you, and it's precisely why it matters. We're rooting for every frontier lab. They're building extraordinary technology, and it makes everything we build better. But our job is to build software that works for you, our users, which means: use tokens productively, don't token-max. This isn't universal advice, if your use case is a customer service chatbot, the calculus is different. But the discipline behind it, spending only what the decision needs, travels everywhere. That's the quest we're on. |
Customer SpotlightQuality Audit Process: Casualty Claim Evaluation Challenge: Resource constraints mean most carriers can only audit 1 to 2% of claims for quality. Solution: QualityLens generates a full quality audit report across 100% of claims, not a sample. Outcome:
Business Impact: Faster audits, more consistent documentation. |
DocLens.ai Product Corner
Feature Spotlight: Automated Claims Quality Audit
Unlike manual quality audits, ClaimLens Quality Audit:
Reads the adjuster's full activity log and every document already indexed in the claim file
Identifies the claim type and automatically skips checks that don't apply
Evaluates the claim against your organization's Expectations Documents, Guidelines, and Regulations
Applies deterministic deadline checks where a missed window is a clear failure
Requires both activity log evidence and supporting documentation before making a judgment-based call
Produces a structured compliance report with evidence behind every finding
Human auditors stay in the loop, reviewing results, overriding findings with written justification, and finalizing the audit. The finalized report becomes a locked, audit-trailed PDF, a defensible record of what was reviewed, what was found, and why.
What's New In The World Of AI
Audits are becoming the next agentic AI frontier
Most agentic AI in claims has lived at intake, systems that read a first notice of loss, verify coverage, and route the claim without a human touching it. The center of gravity is shifting toward the back end of the workflow. Carriers increasingly expect these systems to capture not just the decision but the reasoning behind it, what data informed it, which rules or guidelines applied, and what alternatives were considered, because a judgment call without an audit trail doesn't hold up to a regulator or a courtroom. That's the same principle behind QualityLens: the report has to defend itself, not just summarize the claim.
DocLens.ai In The News
Last month, we published our fourth benchmarking round against Anthropic's frontier models: ClaimLens against Claude Sonnet 4.6 and Claude Opus 4.8. The result was a 40+ point precision gap that, in our own words, "surprised even us." This month, we extended the benchmark to OpenAI's GPT-5.6 Sol and Terra models. The gap didn't move. Purpose-built AI continues to win across the board, against every frontier model we've tested it against.





