Hero photograph for: Demo It Or It Didn't Happen — The Only AI Standard That Matters

Demo It Or It Didn't Happen — The Only AI Standard That Matters

Most AI productivity claims are performance art. Here is the one test that separates real builders from vocabulary tourists.

· 7 min read

Demo It Or It Didn’t Happen — The Only AI Standard That Matters

The Two-Item List

Last month I asked three engineers at three different AEC firms one blunt question: “If I rip out one AI workflow tomorrow, which one actually breaks delivery?”

They paused. Each gave me two things. One said meeting-note summarization. One said spec-sheet formatting. One said RFI triage on a single project. That’s the whole list.

These are the same people posting about agents, RAG pipelines, and autonomous QA. In the real work, the load-bearing AI is two items wide.

That gap — the big talk versus the tiny backbone — is the story. Elena Verna calls it AI Confidence Theater. Call it what it is: bad proof standards. It’s the biggest tax on real builders right now.

The “Show Me” Test Is Not Rude, It Is Hygiene

Engineer conducting a live AI tool demo on a laptop inside a construction site office trailer with project data on screen

Verna’s move is clean and fair. When someone says an AI workflow changed their life, say: show me. Live. Your screen. Your data.

Most “life-changing” claims collapse into Slack summaries, email drafts, and calendar nudges. Useful, sure. Not the revolution the caption promised.

I run vendors through the same drill. You pitch an “autonomous site‑monitoring agent”? Great — point me to one live project where it ran unsupervised for 30 days. Not a sandbox. Not a highlight reel. A real account I can inspect. Half of them disappear. The rest quickly relabel the product as “a copilot to help engineers review site footage.” That’s a fine tool — just not the one they were selling five minutes earlier.

This isn’t hostile. It’s how we treat every other technical claim. If a structural engineer hands you a beam, you ask for the calc sheet. AI is the only place we let vibes count as evidence.

Same Hustle, New Props

Verna nailed the pattern: today’s AI flex is yesterday’s 5 a.m. cold plunge. Same performance. New prop.

The 5 a.m. post said, “I’m more disciplined than you.” The agent post says, “I’m more leveraged than you.” Both chase the same prize — status — and neither has to survive contact with real work.

Once you see it, you can’t unsee it. “I built an agent that writes all my structural calc reports overnight” is doing the same job as “I meditate for 90 minutes before my first meeting.” It’s a status signal disguised as productivity.

The brag isn’t the real problem. The fake baseline is. New engineers start believing everyone else cracked autonomous multi‑agent QA, and their honest two-hour weekly win feels small and shameful. That shame is the damage. It pushes people to perform results instead of compounding real ones.

Hiring Broke First, And Nobody Wants To Say It

Job candidate working through a real RFI dataset on a laptop at a plain desk with a timer and printed documents nearby

Here’s the part that should make every hiring manager sit up: verbal interviews don’t work for AI roles anymore.

Give anyone a week with a good LLM and they’ll place “RAG,” “MCP,” “vector store,” “tool‑calling,” “agentic loop,” and “eval harness” in the right grammatical slots. The words stopped mapping to skill. They map to the same three podcasts.

I’ve watched candidates explain retrieval like professors and then, with two hours and a real dataset, fail to ship a working ingestion script. Not fraud. Just a full decouple between talking and doing.

The fix is boring and unbeatable: work trials. Hand them a real RFI dataset, a folder of messy site photos, or a set of clash reports. Two hours. Build something that runs. You’ll learn more than in five rounds of questions. Anthropic does take‑homes. Shopify runs paid trials. Linear runs project sims. None of this is novel — it’s just now required at every level, not only staff and up.

If your AI hiring loop doesn’t include the candidate touching real data with their own hands, you’re hiring vocabulary.

Why Your FOMO Is Manufactured

Confidence theater sticks because three forces push you there.

Social platforms pay out reach for hyperbole. “AI saves me 15 annoying minutes a week on formatting” gets 40 likes. “I replaced ops with 12 agents” gets 40,000. The feed isn’t neutral. It amplifies hype.

AI moves faster than verification. By the time someone tests a claim from three months ago, the poster is on to the next one. No accountability layer exists.

Inside companies, louder AI claims climb faster. “We automated 30% of QA with agents” earns a slide in the board deck. “We built one reliable summarization workflow that saves three hours a week per PM” earns a polite nod. Guess which one gets funded next quarter.

So that pit in your stomach while you scroll LinkedIn? It’s not a personal flaw. It’s the outcome those systems are designed to create.

What Actually Works, Named Specifically

Monitor in a dim office displaying plain-text AI outputs including RFI triage and meeting summaries for an AEC project

Here’s what a real, non‑theatrical AI stack looks like on a builder’s desk today.

  • Meeting‑note summarization with Notion AI or Granola feeding a searchable project log. It’s plain. It saves about 90 minutes a week. Yank it and the team stumbles.

  • RFI triage with a retrieval layer over past project responses — no magic agent, just semantic search plus a suggested‑answer prompt a human approves. Response time drops ~40%. You can ship it in three days.

  • Spec‑conflict flagging on drawing sets using a scripted OCR pass and a rules library, with an LLM only at the end to explain the conflict in plain English. It’s reliable because the rules carry the load. Think Stripe Radar: rules do the deterministic work, models add last‑mile judgment, humans approve.

  • Firecrawl scraping your own blog into a Lovable app so a domain assistant answers questions grounded in your writing. Verna uses this. It works because the scope is razor‑narrow.

None of these will rack up 40,000 likes. All of them pass the show‑me test. All of them would break real work if you pull them out. That’s the real frontier.

The Standard, For Monday

Pick one AI workflow you actually use this week. Ask two questions.

One: if I remove this tomorrow, do we break, or do I just feel a little less clever? Two: if a teammate asks for a live demo right now, on real data, can I do it?

If both answers are yes, you’ve got a real thing. Name it. Measure the time saved. Share it with the number, not the fairy tale. “This saves me two hours a week on RFI drafts” will earn more respect from serious people than any “changed my life” post.

If either answer is no, that’s theater. Fine — most of us have some. Just stop using it as proof. Don’t let it set your baseline. And don’t let it into procurement or hiring decisions.

Demo it or it didn’t happen. That’s the standard. Everything else is a 5 a.m. cold plunge with a new prop.