Authentica
All resources

AI StrategyAugust 2026

The Sand Castle Is the Strategy

AI has made the prototype nearly free: a working model of a business process in a day, a convincing demo in a week. Most enterprises then make one of two mistakes. They promote the demo to production, or they throw it away and pay for the discovery twice. The third path is to treat every prototype as a survey of the ground, and to keep what it finds.

A team traces a reusable blueprint from a sandcastle before the tide arrives.

The analogy comes from a recent conversation between Benchmark’s Eric Vishria and Patrick O’Shaughnessy [1]: the software being built in the AI era is sand castles, not stone, and the winners will be the ones who are comfortable with that. Most enterprises hear the analogy as a warning. I think it is the strategy, provided you are disciplined about what the tide is not allowed to take.

The fastest prototypes in the history of software

We have never been able to build software this quickly. You can sit down with someone who understands a business process, talk for an hour, and by the end of the day have something that looks remarkably close to the finished product. It reads their documents, speaks their vocabulary, queries their systems, reasons through an exception, produces an analysis, and, if you let it, takes an action. Work that would have been a two-quarter integration project in 2020 is now an afternoon with a capable model and a person who knows where the bodies are buried.

It is also a sand castle, and that is not an insult. The prompt holds the process together. The agent framework is whatever the builder reached for that week. The rules everyone considers obvious live in a system prompt that nobody has reviewed. None of it would survive an audit, a model deprecation, or the departure of the one person who built it. A sand castle is exactly what it is, and a sand castle is exactly what you should be building right now.

The traditional enterprise instinct is the opposite: start by pouring concrete. Architecture committees, requirements documents, integration specifications, a six-month implementation plan. That instinct was rational when the underlying technology moved slowly enough that a two-year-old architectural decision was still a good one. AI does not extend that courtesy. Models improve every few months. Agent runtimes appear and disappear; a technique that counted as sophisticated engineering a year ago is now a primitive exposed by an API. In an environment moving that fast, committing to a permanent architecture before you have learned what is actually useful is a reliable way to spend a great deal of money building the wrong thing.

A sand castle, by contrast, is cheap to change. You can knock down a tower and rebuild it ten feet away. You can discover that the feature everyone insisted was essential does not matter, while the ugly little structure someone improvised in a corner turns out to be where the value is. Vishria makes a related observation about the companies themselves: the competitive frontier has shifted, everything is, in the episode’s phrase, a jump ball, and the advantage goes to the teams that iterate along the jagged edge of model capability faster than the edge moves [1]. His exemplar is Cursor, which has obsoleted its own six-month-old product repeatedly, an IDE, then tab completion, then agentic work, each version demolishing the one its users had just adopted. The pattern worth copying is not resilience to disruption. It is self-obsolescence: the winners are not the castles that withstand the tide, they are the tide. An enterprise adopting AI is in the same position as an AI company defending one. Build quickly, put it in front of the people who do the work, and learn.

This is not a posture we recommend from the sidelines. In February I argued that vendor selection was about to reduce to what investors call the prediction premium: the outsized credibility that accrues to bets that diverge from consensus and prove right. The bet we offered as evidence was one we had placed a year earlier, that AI tools built for software engineers would become tools for all knowledge work, which is why we were already running our own back office through Claude Code when Anthropic’s Cowork and OpenAI’s Frontier arrived, within a month of that essay, and made “agents as coworkers” the industry’s consensus. We did not see it early because we are better forecasters. We saw it early because we were building sand castles on our own operations, and the frontier is legible from up close in a way it never is from a conference room. The prediction premium is the return; the sand castles are how you earn it.

The problem begins when the tide comes in.

When the tide comes in, most strategies allow two outcomes

The first is to pretend the sand castle is a real castle. The demo works, so it quietly becomes production. The prompt becomes the business logic. The agent framework becomes the architecture. A collection of tool calls becomes the controls environment. The prototype accumulates integrations, credentials, exceptions, and production dependencies until replacing any part of it feels dangerous, and the organization is now operating critical workflows on a structure that was never designed to bear weight. It works until it doesn’t, and when it doesn’t, nobody can say precisely why, because the rules it was enforcing were never written down anywhere a reviewer could read.

The second outcome is almost as expensive: throw the sand castle away and build the real thing from scratch. The enterprise discovers that the prototype was cheap but production is a separate project. The workflows are reimplemented, the integrations rebuilt, the rules translated from a prompt into code, the tests recreated from memory, and the trust of the operators, who reasonably believed the thing already worked, re-earned from zero. The organization pays twice: once to discover what it wants, and again to build it.

The second payment is made in a scarcer currency than money. Most enterprises are not short on AI ambition; they are short on capacity. Every pilot consumes months of vendor evaluations, consultant engagements, and internal alignment before a single use case gets tested, and by the time the proof of concept produces results, the team that sponsored it has moved on to firefighting something else. The pattern repeats: another vendor, another discovery phase, another deck. Adoption stalls not because the technology fails but because the process of evaluating it exhausts the organization.

This failure mode has been measured. MIT’s 2025 State of AI in Business study, the largest of its kind, found that roughly 95% of enterprise generative-AI pilots produced no measurable P&L impact, across some $30 to 40 billion of investment [2][3]. The study’s diagnosis is more specific than “the models are weak” or “the use cases are wrong.” The authors call it a learning gap. In the report’s words, “the core barrier to scaling is not infrastructure, regulation, or talent. It is learning. Most GenAI systems do not retain feedback, adapt to context, or improve over time” [2]. The demo proves the capability is possible and then leaves nothing durable enough to build on. Every lesson the pilot surfaced evaporates with the pilot.

There is a third path, and it is the one this piece argues for: the sand castle should become the blueprint for the real castle.

What has to survive the prototype

Taylor’s version of the analogy, which Vishria endorses, holds that the washing away is fine. “We used to be building castles,” Vishria says, quoting Taylor. “Now we’re building sand castles that are going to get washed away,” and the artisan who wants a perfect foundation to stand for a hundred years “is just not going to make it” [1]. I agree, with a qualification that determines what an enterprise should actually do: things should wash away in proportion to what they cost to rebuild. For an AI company iterating on its own product, nearly everything is cheap to rebuild, so nearly everything can be sand. The remarkable fact about this moment is that code, which for seventy years was the expensive artifact, the thing enterprise architecture existed to protect, has moved to the cheap side of the ledger. This migration was visible before it was consensus; more than two years ago I wrote that the cost of software was headed to zero, and the ledger has been rebalancing ever since [4]. What sits on the expensive side now is semantics: what things mean in this organization, how the process actually runs as opposed to how the SOP says it runs, who holds which authority. The durable asset is therefore not the things that encode the process, the code, the prompts, the agent graph. It is the formalization of the process itself. Code can be sand precisely because meaning is stone.

When we talk about optionality at Authentica, this is what we mean, stated as an inventory. The model does not need to survive the prototype. Neither does the agent framework, the infrastructure provider, or most of the code. What has to survive is what you learned about the business, because that is the part that was expensive to acquire and that no vendor can sell you.

Concretely: what an invoice is in this organization, as opposed to in the textbook. What constitutes a discrepancy, and which discrepancies resolve automatically versus requiring approval. Who can approve, and what evidence must exist before they do. Which systems contain the facts and which fields in them cannot be trusted. What “done” means. Which situations have already made the system fail. Every one of those is a discovery, most of them are undocumented before the prototype forces them into the open, and together they are the stones the real castle is built from.

So we capture them separately from whatever happens to be doing the reasoning that week. At Authentica the vessel is the customer’s operating model: a typed, versioned description of the entities, actions, tasks, policies, and authority structure the agents operate against, composed independently of the model or runtime underneath it. In the newest generation of the platform this extraction has its own surface, Studio, where the operating model is authored and versioned by your experts and our agents together, while the prototype is still standing. Experimentation stops being disposable. Every sand castle leaves something behind.

What the prototype surfaces becomes what the enterprise owns
The left column is discovered in weeks and lives nowhere durable. Extracted into the operating model, each discovery becomes a versioned artifact that outlives the prototype that found it.
Surfaced by the prototype — weeks Owned in the operating model — years What an invoice means here not what the textbook says it means Entity in the ontology The action the agent takes currently implicit in a tool call Explicit, typed action contract The threshold that needs a second review currently a sentence in a prompt Policy, versioned and enforced The exception that looked wrong but wasn't institutional knowledge, previously oral Evaluation case The correction a person made today it vanishes into a chat log Gold case in the benchmark The ERP field nobody trusts every veteran knows; no system records it Capability interface The prototype is disposable. The description is not.
The mapping is the discipline. None of it happens automatically; each row is a deliberate act of extraction, done while the prototype is still standing.

Sand becomes stone one layer at a time

A worked example makes the mechanics concrete. Take a freight-audit workflow. On day one, the right system to build is a deliberately loose one: give an agent a set of invoices, let it analyze them, keep a person in front of every decision, and use the strongest model available, because at this stage you are discovering the shape of the problem, not optimizing its economics. That is a sand castle, and deliberately so.

While people use it, the discoveries arrive. A charge with this code means one thing for this carrier and something different for another. An invoice above this threshold needs a second review. This exception looks suspicious and is in fact routine. This field in the ERP cannot be trusted. This person is the real approver, whatever the SOP says. Each of those is extracted as it appears, into the artifacts of the figure above: entities into the ontology, actions into contracts, thresholds into policy, the strange edge case into an evaluation, the human correction into another gold case, the untrustworthy field into a capability interface that says so.

Then the system is tightened, one axis at a time, against the same description. The data moves from synthetic to live. The access moves from read-only, to proposing changes, to taking narrowly governed actions. The approval posture moves from every action requiring sign-off to only the uncertain cases escalating. And the model, which began as the most capable one money could rent, becomes whichever one passes the benchmark you have been accumulating since day one, a substitution we have measured directly: under run-time enforcement, a cheap open-weight model matched a frontier model on real compliance work. The intelligence underneath changes throughout the tightening. The operating model does not have to.

The tightening: every axis moves except the one you own
From first prototype to production, each operational axis is tightened independently. The operating model is the constant they are all tightened against.
Day one Pilot Production Scale Data synthetic live Access read-only proposes changes takes governed actions Approval a person in front of every action only uncertain cases escalate Model the strongest available, uneconomically the cheapest that passes your benchmark The operating model — versioned, owned, unchanged in kind every stage above is tightened against the same description, so autonomy is granted by policy, not by rebuild
The boundaries between stages are governance decisions, not engineering projects. Loosening an axis back down, after an incident, say, is a policy change with an audit trail, not a rollback of the system.

There is an inversion hiding in this sequence worth making explicit. The exhausted enterprise slows its business down to evaluate AI. Run this way, the AI speeds up how the business designs and tests what it needs: because the agents do most of the construction, working inside guardrails already set and against a model of how the business actually runs, a use case goes from idea to working agent in weeks rather than quarters, and the evaluation cycle that produces the fatigue never starts. The deployments also compound. What one agent learns about the operation, extracted into the operating model, becomes the foundation for the next, and often surfaces opportunities nobody had scoped. Adoption stops being a series of exhausting one-off projects and starts behaving like momentum.

Optionality is not indecision

There is a version of “vendor optionality” that is really a refusal to decide: never commit to anything, abstract every component, build for twelve hypothetical futures. That produces terrible software, and it is not what I am arguing for. You should use the best tools available today, by name. Use OpenClaw if it is the right runtime. Use Claude if it is the best model for the job, and OpenAI next month if it becomes better. Connect through whichever integration platform gets you live fastest. Pick things and move.

The architectural question is narrower and more important: which decisions become embedded in the thing you own, and which remain replaceable implementation choices? Our answer is that the business semantics belong to the customer, the ontology belongs to the customer, the benchmark that defines “correct” belongs to the customer, and the workflows and software built on their behalf belong to the customer, while the model, the runtime, and the infrastructure stay replaceable. We have written elsewhere about why reliability follows from separating the operating model from the model and why one operating model can govern agents across different runtimes. The sand-castle argument supplies the reason that separation matters during implementation, not just in the steady state: it is what preserves the learning.

The real lock-in is losing the lessons

Lock-in is usually discussed in terms of contracts and data export, and those matter. But AI creates a more dangerous form that neither procurement clause addresses: knowledge lock-in. I have felt it in miniature. In 2024 a bug erased the memory my ChatGPT account had accumulated, a year of taught context gone in an afternoon, and I wrote at the time that agent state needed version control, a problem every major AI platform now treats as live [4]. The enterprise version is the same event at scale. An organization spends a year teaching a system how it really works. Thousands of corrections. Hundreds of edge cases. The exceptions nobody documented, the approvals that operate differently in practice than on paper, the definition of quality a twenty-year veteran carries but has never written down. If all of that ends up encoded in one vendor’s prompts, proprietary agent graph, fine-tuned weights, or application code, then switching providers means losing part of the organization’s own self-knowledge, the part it paid a year of corrections to externalize. Your data can be perfectly exportable and you can still be trapped.

The alternative is to make the lessons first-class, customer-owned artifacts, which is why I have come to think of the ontology less as a feature of an AI deployment and more as the asset the deployment produces. The agent is doing work, and that work pays for the project. But underneath it, the enterprise is gradually acquiring a machine-readable description of itself, one that gets more accurate every time reality corrects it. The MIT study’s learning gap is precisely the absence of this asset: pilots that cannot retain what they are taught [2]. A pilot that writes its lessons into an operating model cannot help but retain them.

A castle you can still move

The analogy does break in one place, and better to name it than stretch the metaphor. Real castles are hard to move, and their builders considered that a feature. Enterprise AI cannot afford to. The failure state on the far side of prototype chaos is a stone structure so safe it cannot change: an architecture that survives audits and outlives its own usefulness. The goal is a durable foundation with replaceable machinery on top, a production system that can state its rules, enumerate its actions, bound each agent’s authority, measure its own correctness, and produce evidence of what happened, and then run all of that on whatever intelligence and infrastructure make sense this quarter.

That is a different definition of enterprise-grade than the one the industry inherited. Enterprise-grade used to mean built not to change. In AI it increasingly means built so that change does not destroy what you have already learned.

Build sand castles aggressively

So the advice to enterprises is not to slow down; it is close to the opposite. Build more sand castles. Experiment with the newest models. Try runtimes that did not exist last quarter. Automate the process everyone insists is impossible, and put the prototype in front of operators before the architecture committee has scheduled its second meeting. Then demolish your own pilots on purpose, the way Cursor demolishes its own product, because self-obsolescence is only expensive when the lessons are stored in the thing being demolished, and by now yours are not. That 95% is what speed costs when nothing survives it [2].

The discipline is entirely in what happens next, and it fits in five sentences. When you discover something true about your business, extract it from the prototype. When a person corrects the agent, turn the correction into a benchmark case. When you find a rule, encode it somewhere you own. When a workflow proves valuable, separate what it means from whichever technology currently executes it. When you grant an agent more autonomy, grant it against the same operating model rather than rebuilding the system around a more powerful agent.

Do that, and the speed of the sand castle stops being technical debt and becomes the survey instrument it should have been all along: the mechanism through which you discover, cheaply and in weeks, where the real castle’s foundations belong. Move at prototype speed without accepting prototype permanence. Use sand to find the shape. Keep the blueprint. And when the tide comes in, as it will, make sure the most valuable thing you built is still standing.


References

  1. “Sandcastles & Silicon.” Eric Vishria with Patrick O’Shaughnessy, Invest Like the Best, 2026. https://colossus.com/episode/sandcastles-and-silicon/ — Vishria credits the sand-castle framing to Sierra’s Bret Taylor.
  2. “The GenAI Divide: State of AI in Business 2025.” Challapally, Pease, Raskar, and Chari, MIT NANDA, July 2025. Report PDF: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  3. “MIT report: 95% of generative AI pilots at companies are failing.” Fortune, August 2025. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
  4. Prediction record: multi-agent experiments pre-reasoning-models, early enterprise RAG, version control for agent state, and the cost of software going to zero. Michael Borg, LinkedIn, February 2026. https://www.linkedin.com/posts/activity-7422267197235843072-Gs5I

Related: Intelligence Sovereignty: Owning Your Alpha in a World of Rented Minds · The Model Is Not the Product: Why Reliable AI Runs on an Ontology · The Sisterhood of the Travelling Ontology · Why We’re Building OrgBench™

If you’re deciding what should outlive your next prototype, reach out for a conversation.