pq
Free Playbook · Kickoff → Production

Our entire 90-day playbook. Free.

Perpetual Quest · 2026

The 90-Day
Deployment
Playbook

A 90-day path to production. Go-live only when the evidence gates are met.

Qualify and blueprint, controlled prototype, shadow test, human-approved production, bounded autonomy, and stabilization — the exact structure we use to take an AI system from kickoff to production. Use it with us or without us.

Instant access · Also sent to your email · No spam

Dear Operator,

This is the entire method. Not a teaser, not chapter one of a sales funnel — the actual phase-by-phase structure we use on a 90-day path from kickoff toward production. The calendar is a target; go-live happens only when the production-readiness gates below are met. We’re giving it away because we deploy, we don’t describe — and because operators respect the real method, and most won’t execute it alone anyway.

One framing rule before the phases: the goal is not a prototype, a pilot, or a proof of concept that dies in a sandbox. The goal is a system your team uses every day, owns outright, and can operate and extend without us. Choosing Run later for managed assurance — evaluation, governance, optimization, incident support — is a choice, not proof of a failed handoff. Every decision below serves that.

The six phases

Days 1–15

Qualify and blueprint

We map the target workflow end-to-end: who touches it, what systems it crosses, where the hours go, and where judgment is genuinely required versus where it’s habit. Nothing gets built until the scope and success criteria are explicit.

Three artifacts come out of this phase, and all three matter:

  • A written scope — one workflow, defined boundaries, and an explicit “not in this deployment” list. Scope creep is the #1 killer; this document is the guardrail.
  • A success metric — the number we’ll be judged on at day 90: hours saved, cycle time, error rate. Agreed before a line of configuration.
  • A named internal champion with reserved bandwidth — 4–6 hours a week, on their calendar, approved by their manager. We’ve watched deployments die because the champion had no bandwidth, so we settle this before we scope anything else.

Days 16–30

Build the controlled prototype

The first build is intentionally controlled. Integrations, data plumbing, and the agent workflow come together in an environment where we can trace decisions and prove the basics before any write authority is allowed.

  • Retrieval and data plumbing get proved first. The unglamorous part determines whether everything else works.
  • The agent workflow runs against historical and sandboxed work with structured outputs and visible traces.
  • Escalation routes and edge-case handling start here, not as an afterthought after launch.

Days 31–45

Shadow test

Now the system runs in parallel with real work. Humans still do the actual job while the agent shadows the workflow so we can compare judgment, catch disagreement patterns, and learn where the edge cases live before anything customer-facing changes.

Scope-change requests during shadow go through one filter: anything added pushes something out. Written down, every time. That discipline is what gets you to production while everyone else’s pilot is still “evolving.”

Days 46–60

Human-approved production

The system goes live only after the production-ready evidence gates below are met. At launch, a human review loop covers every high-stakes output at first, then progressively samples as trust is earned with evidence.

  • Customer-facing or externally consequential actions stay approval-gated.
  • Decision quality and execution reliability have to show up on live work, not just in staging.
  • The success metric from blueprint is tracked publicly, so everyone can see whether it’s working — including when it isn’t yet.

Days 61–75

Bounded autonomy

Once the evidence says the system can be trusted, selected low-risk actions can run inside explicit limits. The rule stays the same: agents handle volume, humans handle exceptions, and the exception paths get built with the same care as the happy path.

This is where volume reveals the edge cases staging never will. We sample output, tune quickly, and keep the authority bounded to the cases that have actually earned it.

Days 76–90

Stabilize and transition

The final phase is stabilization and transition. Daily monitoring, fast tuning cycles, runbooks, admin access, and named owners all have to be in place before the deployment is done.

The exit test is simple: if we disappeared on day 91, the system keeps running and improving. Most consultancies structure engagements so you need them forever because the system only works with them. We structure them so you can run without us — and continue with Run only if ongoing evaluation, governance, optimization, and carefully managed expansion are worth it to you.

Training runs across all six phases

Training is not a phase that happens after launch. It runs alongside the deployment so the operators who will own the system are learning on real work, not demo data. Teams trained after the fact treat the system as something done to them; teams trained during rollout treat it as something they own.

  • Weekly working sessions with the people who’ll actually operate the system.
  • The champion learns configuration, not just usage: how to adjust prompts, thresholds, and routing. Material prompt, tool, or permission changes after go-live follow the Agent Change Control Standard (especially under Run) — not silent production edits.
  • By launch, the team has already used the system for weeks. Day one of production is nobody’s first day.

What “production-ready” means before go-live

Going live is not the definition of production-ready. Before we call a system production-ready, these evidence gates are met. Thresholds for the metric gates are agreed in scoping week — we do not invent universal pass rates; we define them with you, then prove them.

  • Evaluation pass rate — agreed threshold met on a fixed eval set before go-live.
  • Critical failure rate — within the agreed bound on shadow/parallel runs.
  • Escalation accuracy — exceptions route correctly at the agreed rate.
  • Tool-call accuracy — tool use succeeds within agreed bounds.
  • Maximum retry limits — hard caps documented and enforced.
  • Rollback mechanism — documented, tested path to last-known-good.
  • Kill switch — named owners and time-to-pause (see the governance checklist).
  • Audit completeness — reconstructible action history for the workflow (same checklist).
  • Cost-per-completed-workflow — baseline measured; within the agreed envelope.
  • Performance under realistic volume — validated at expected peak, not demo load.
  • Named owner for every failure class — map of failure modes to accountable humans.

Why 90 days is a target (and when go-live waits)

The honest framing: 90 days is the path, not an unconditional launch date. Target production deployment within 90 days, subject to the agreed production-readiness gates. If day 90 arrives and a critical failure rate (or any other gate) is still above the threshold you agreed in scoping, we do not go live. The evidence standard wins over the calendar.

The other things that break the timeline are scope creep (solved by the written scope and the trade rule) and discovering mid-build that the underlying process is broken. AI doesn’t fix broken operations — it accelerates working ones. If scoping week reveals chaos underneath, the right move is a process redesign first, and we’ll say so directly rather than deploy on sand.

What this produced in our own businesses

This structure isn’t theory — it’s how we run our own companies. It’s the method behind Grademate scaling enrollment 3× with zero added delivery headcount, NamingForce compressing a review cycle from hours to under 20 minutes per project, and EchoTexting absorbing a workload that would otherwise have required two to three additional ops hires. We broke it ourselves first. That’s the point.

Use it — with us or without us

Everything above is enough to run a deployment yourself if you have the internal muscle. If you’d rather have the people who’ve run this loop dozens of times do it with you, that’s a Forge engagement: a 90-day path to one system in production, your team trained and owning it — with go-live only when the evidence gates are met.

Book the 30-minute call — we’ll map your first deployment →

— Perpetual Quest
perpetualquest.com · [email protected]