Adoption with guardrails — built by a security team

Your Team Has AI. Nobody Taught Them How to Use It.

Training, governance, and cost control for companies that bought the licenses and are still waiting for the payoff. We do adoption the way we do security — practical, right-sized, and honest about what it can't do.

40% of desk workers received AI “workslop” last month — output that looks like work and isn't. Two hours to fix. Every time. The cost of untrained AI isn't the license. It's the rework it pushes onto everyone downstream. — BetterUp Labs & Stanford Social Media Lab, n=1,150, September 2025
Artifacts you own — not a slide deck
Every deliverable is built from your work, not from a template. Click any card for detail.

A Prompt & Skill Library From Your Actual Work

  • Not generic examples — the five things each role does every week
  • Tested before it goes in
  • Yours to keep and extend

A Model Routing Policy That Cuts the Bill

  • Which work goes cheap, which goes expensive
  • A rule your people can actually remember
  • Benchmarked on your tasks, not a leaderboard

A Context System That Survives the Session Ending

  • Project instructions, handoff templates, a save routine
  • Nobody re-explains the same thing every morning
  • With a size budget, so it doesn't eat your context

A Skill Catalog With Owners, Ratings, and an Expiry Date

  • Keeps the library from rotting into 400 near-duplicates
  • Rating plus usage, so the good one wins the search
  • Nothing lives forever unreviewed

An AI Policy People Can Actually Follow

  • Data classification, approved tools, what never goes in a prompt
  • Enforcement gaps named honestly, in the document
  • We write these for a living

A Shadow-AI Inventory

  • What your people are already pasting into what
  • Usually uncomfortable
  • Always the most useful page in the readout

Four numbers that shape how we run these engagements. Every one has a primary source — click for it.

40%
of desk workers got AI
“workslop” last month
60%
cheaper to run the same task
one model tier down
−24pts
how much worse trained users did
outside the model's competence
+42.5%
quality gain when they
were inside it

The bill isn't the cost. The cost is the rework nobody invoiced you for.

Training alone can make it worse

758 consultants. Three groups: no AI, AI, and AI plus prompt training.

+42.5%

On work the model was good at, the trained group won. Better quality than untrained AI users. 93% of tasks finished, against 82% for the group with no AI at all.

−24 pts

On work the model was bad at, the trained group was wrong 24 points more often than the people with no AI — and nearly twice as wrong as untrained AI users. Their answers were also better written. Polished, confident, wrong.

Why? The researchers found the training increased copy-and-paste. More skill, more trust, less checking.

Prompt training alone makes people faster, more productive, and more confidently wrong.

So we don't teach prompting alone. Every HyperX engagement teaches the same person how to prompt and how to tell when an answer is out of the model's depth. That isn't a module we added to fill a day. It's the reason the rest of it works.

Dell'Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon & Lakhani. “Navigating the Jagged Technological Frontier.” Organization Science 37(2), 2026. Pre-registered field experiment with Boston Consulting Group.

Everyone sells the training. Almost nobody mentions the failure mode it creates.

Three tracks — choose your depth
Start anywhere. Most companies start with the assessment, then pick a track from what it finds.
Right-sizing, out loud: a 30-person company doesn't need a program. It needs Essentials, a policy, and a routing rule. Two weeks, done. We'll tell you that before you buy it.
Every track ends with artifacts built from your work, not ours. The exercises use your real tasks, the skills get built for your team, and the policies name your tools. That's the part a syllabus can't hand you.
Most of your AI bill is the wrong model doing easy work
Three beats, and then the arithmetic.

The premium is real

The top-tier model costs a multiple of the mid tier. That premium buys real capability — hard reasoning, long complex code, ambiguous judgment.

That isn't what most work is

Summarize this thread. Reformat this into a table. Draft a reply. Classify these tickets. On that work the cheaper tier is indistinguishable.

Nobody routes

Because nobody told them to. The default is whatever the button is set to, and the button is usually set to the most expensive option.

The arithmetic

Published list pricing, August 2026. Ratios, not dollars — vendors re-price.

TierCost vs. top tierSavings
Topbaseline
Mid40%60% cheaper
Small20%80% cheaper

Move one workload down one tier and you cut its cost 60%. Move it two and you cut it 80%. Nothing about the work changed. Only the routing did.

This is why we quote 40–60% and call it conservative. It isn't our estimate — it's arithmetic on published prices. A team routing everything to the top tier and shifting its volume work down saves more than 60%. We say 40–60% because a real blend keeps the hard work where it belongs.

The routing rule

The actual deliverable is one page your people can remember.

The workRoute toWhy
Summarize, reformat, extract, classifySmall tierIndistinguishable quality at 20% of the cost. Volume work lives here.
Drafting, analysis, meeting prep, first draftsMid tierThe workhorse. Most seats should default here.
Hard reasoning, long code, ambiguous judgment, high-stakes reviewTop tierWorth every cent — for the 10–20% that needs it.
Anything in a loop, on a schedule, or over thousands of recordsCheapest tier that passes the testAutomation is where the wrong default gets expensive fastest.

Honest caveat: savings depend on your current mix. A team already on the mid tier won't see 60%.

The table above is the easy part. Knowing which of your workflows belong in which row takes a benchmark on your actual tasks — not a leaderboard, and not a guess. The assessment tells you your number before you commit to anything.

You won't see it happen. You'll see the mistake it makes ten minutes later.

The model forgets. Your people won't notice.
Every session has a finite window. When it fills, the system doesn't crash — it summarizes older messages away and keeps going.

Summaries are lossy, and the loss is not random. The rule survives. The reason for the rule doesn't. The constraint survives. The exception you carved out doesn't.

What that actually looks like

The half-finished plan

A twelve-step deployment plan, compacted at step seven. The model completes step eight and declares the project finished — dropping verification, rollback prep, and notification. It didn't say it was skipping them. It said “done.” It thought it was.

The TODO that shipped

A config file written with placeholders and a plan to fill them in later. Compaction killed the plan, not the file. The model saw a clean file and committed it. The next service failed at startup parsing the literal placeholder.

The exception that dissolved

“Never edit this directly” became “be careful with system files” in the summary. And “be careful with” is, to a confident model, an invitation to do the thing carefully rather than not at all.

The model doesn't know what it forgot. You can't ask it to check a constraint it no longer remembers having.

Two things go wrong, and they're not the same

This distinction is the whole lesson. Almost nobody has a name for the second one.

FullnessDepth
What it measureshow close you are to the summarize-and-continue pointhow much context you're carrying
The riskyou lose contextthe model gets duller — misses details, repeats work, drifts
NatureHard. A wall.Soft. A gradient.
The fixa deliberate checkpoint, or the recovery playbooka clean checkpoint, whenever it suits you
UrgencyNowNot urgent

The window limit isn't the quality limit.

And bigger windows made this worse, not better. On a million-token window you can be deep into the degradation range at 20% full — because 20% of a million is still 200,000 tokens of accumulated context. The bigger window didn't fix the problem. It removed the wall that used to force a reset.

What we teach

Detection

It re-asks something already answered. Suggests an approach you ruled out. States a fact differently than you remember stating it. The “wait, why is it suggesting that?” moment. When you spot one: stop. Treat the last ten minutes as suspect.

Recovery

Stop and verify. Re-read the source docs — don't paraphrase. Check against ground truth, because the file on disk beats the summary in the model's head. Re-state the constraints, since they're exactly what got softened. Resume from a known-good point, not from where the model thinks you were.

Prevention

Scope work to finish before the wall. Check before you start something big. And keep your always-loaded context files small — every line in them is paid for in every session, forever.

The gauge

Two indicators. That's the whole tool.

Hard · allowed to be loud

% Until Compact

How close you are to the wall — measured against the point where summarizing actually kicks in, not against the size of the window. This is the one with an alarm on it.

Soft · a nudge, never an alarm

Context Depth

How much you're carrying. Not urgent, ever — just a signal that a clean checkpoint would sharpen things up.

🟢 Sharp🟡 Warming🟠 Drifting🔴 Deep

And one thing we'll tell you that nobody else will

We built that depth meter with a red danger threshold. Then we measured it against real session history, looking for the point where quality falls off.

There wasn't one. The error rate came out flat.

So we softened the meter so it never reaches alarm red, and we put a card in the tool that says “measured from your history — no knee.” It's a research-informed heuristic, deliberately gentle, and we label it that way. We could have kept the scarier number. It would have sold better.

We built a warning light, checked our own data, found the danger wasn't where we'd drawn it, and shipped the correction.

That's the same instinct behind everything else on this page.

The thresholds are the easy part. Knowing which of your workflows actually run long enough to hit them — and which of your people are quietly working through it every day — is what the assessment finds.
400 skills, no catalog, and nobody trusts any of them

Year one, everyone's thrilled. People build custom skills, agents, saved prompts. Year two there are hundreds. Twelve do nearly the same thing. Four are subtly broken. The person who wrote the good one left. Nobody can find anything, so everyone writes their own — which makes it worse.

This isn't new. It's policy sprawl, script sprawl, and the shared drive with nine versions of the same spreadsheet. We've fixed that before. The fix is governance, not enthusiasm.

Sprawl isn't too many skills. It's too many skills of unknown quality.

What we stand up

A catalog with a required shape

Name, one-line description, owner, what it's for, what it's not for, last-reviewed date, worked example. No metadata, no entry.

Ratings and usage — let the org sort it

A rating after use, a usage count (the honest signal — what people run beats what people praise), and a one-click worked / didn't work. Search defaults to a blend, so the good one wins.

Tiers, so trust is visible at a glance

🥇 Golden — reviewed, owned, tested. Use this one. ✅ Community — someone's working skill, ratings visible. 🧪 Experimental — no promises. 🗄️ Archived — superseded, hidden from search. Plus a promotion path with a named reviewer, so quality is a process rather than an opinion.

An expiry date on everything

Unreviewed two quarters → flagged. Unused two quarters → archived. Sprawl is what happens when nothing is allowed to die.

Duplicate detection at intake

Before you publish, the catalog shows the three closest existing skills and asks whether you'd rather improve one. Most sprawl is honest duplication by people who couldn't find the original.

It lives where they already work

A repo, a SharePoint list, a Notion database, a board. Not a twelfth destination nobody opens.

Without a catalogWith one
400 skills, unknown quality40 golden, the rest tiered and searchable
Everyone rebuilds the same thingDuplicate check at intake
Author left, skill rotsEvery skill has a named owner
“Which one works?”Rating + usage answers it at a glance
Nothing is retiredExpiry and auto-archive
The catalog design isn't the hard part. The first pass is — going through what you already have and deciding which ones are golden, which are duplicates, and who owns each survivor. That's a week of somebody's judgment, and it's the week that makes the rest work.

The class is the easy part. Making it survive contact with a Tuesday is the work.

Four steps, and you can stop after the first

Start with the assessment.

Two weeks. Scoped and priced up front, no commitment past it. You'll get:

  • What AI your people are actually using — including what you didn't approve
  • What you're spending, and what routing would save you
  • Your prompt and skill inventory, and its real condition
  • A right-sized plan: what you need now, what can wait, what you don't need at all
Book an AI readiness assessment

We publish the tools we teach with

Everything on this page is public. The methods are in our articles. The tools are on GitHub. What we sell is knowing which of it applies to you — and that part starts with looking at what you actually have.

/adversarial-review

Independent reviewers attack a deliverable before it ships, because same-session self-review is blind to its own gaps.

/sanitize-review

Catches sensitive detail before content leaves the machine — including the dangerous case: harmless facts that together identify one organization.

/save-session /save-full

Structured session records that survive a context reset and hand off between people intact.

/prune

Shrinks a bloated always-loaded context file by relocating what drifted into it — measured in bytes, because line counts hide the problem.

Surviving Compaction · Why I Treat My AI Context Like Infrastructure · Every Rule in My CLAUDE.md Is a Mistake I Made Once · I Taught an AI to Catch Its Own Lies. It Lied During the Lesson. · So You're Good at AI? Build This Skill Before You Get Sued · My AI Hygiene Rule Passed Its Own Check For Months. The Check Was Blind.