AI Implementation Consulting
The part everyone underestimates. We take AI from a promising demo to a system that runs every day, integrated with your stack, monitored, evaluated, and owned by your team.
The Last 20 Per Cent Is 80 Per Cent of the Work
Every AI pilot works. That is what makes them so dangerous as a basis for planning.
Building a demo that impresses an executive team takes a competent developer about three weeks. Building a system your staff will still be using in a year takes considerably longer, and almost none of the additional time goes into the part that was demonstrated. It goes into the invoice photographed at an angle in a ute at dusk. The customer who asks three questions in one message and contradicts themselves in the third. The integration that silently rate-limits on the last business day of the month, which is exactly when the volume peaks.
This is why so many organisations have a graveyard of pilots and nothing in production. The pilot proved the idea was possible, everybody got excited, and then the project met error handling, evaluation, permissions, audit logging, cost control and the human escalation path — the unglamorous scaffolding that turns a demonstration into an operational system. Budget and enthusiasm tend to run out at roughly the same moment, and the initiative quietly dies at 80 per cent complete.
We scope for the whole distance. Our implementations assume the messy input is the normal input, because in your business it is. We build the evaluation harness before the system meets a customer, so “is it working?” has a numeric answer measured on your data rather than a vendor’s benchmark. We design the escalation path first, because the question is never whether AI will hit something it cannot handle — it is whether a human finds out gracefully or the customer does.
And we build for handover from the first commit. Your repository, your cloud accounts, your billing, your runbook. The measure of a good implementation is not that it works on the day we present it. It is that six months later your team has changed it three times without calling us.
What Production Actually Requires
The scaffolding that separates a system from a demonstration. All of it is in scope.
Real Integration
Connected to your systems of record — the ERP, the CRM, the vertical package with the difficult interface. Not a spreadsheet export that someone has to remember to run every Monday.
- System-of-record read and write paths
- Fallbacks for systems without a usable API
- Idempotency, so a retry never double-books or double-charges
- Auth and permissions that respect your existing roles
Evaluation Harness
A scored test set built from your real historical cases, including the ugly ones. Every change is measured against it, so quality is a number rather than a feeling.
- Test set drawn from your actual data
- Accuracy measured on your cases, not a benchmark
- Regression testing on every change
- Explicit pass thresholds agreed before launch
Escalation & Failure Paths
Designed first, not bolted on. When the AI is unsure, a human finds out — with the full context and before the customer does.
- Confidence thresholds that trigger human review
- Full context handed to the person picking it up
- Graceful degradation when a dependency is down
- No silent failures, ever
Monitoring & Cost Control
Every interaction logged. Alerts on the things that matter: escalation rate, latency, accuracy drift, and cost per transaction — which has an unpleasant habit of climbing quietly.
- Full interaction logging and sampling for review
- Alerting on drift, latency and error rates
- Per-transaction cost tracking with budget alarms
- Dashboards your team reads, not ones we read
Your Code, Your Infrastructure
The repository is yours from the first commit. Infrastructure runs in your cloud accounts on your billing. No black boxes, no hostage architecture.
- Code in your repository from day one
- Runs in your cloud accounts, billed to you directly
- Model access behind an interface, so it can be swapped
- No proprietary runtime you have to keep renting
Handover That Sticks
A written runbook, training for the people who will own it, and a period of supported operation where your team drives and we sit behind them.
- Runbook covering operation, failure modes and changes
- Hands-on training for the internal owner
- Supported-operation period before we step back
- Optional light retainer — never a requirement
How We Get to Production
Six to ten weeks for a well-scoped workflow. We do not quote two-week AI builds, because a two-week AI build is a demo.
Scope & Success Criteria
We agree what the system must do and — more importantly — what accuracy counts as good enough to launch. That number gets written down before anyone builds anything, so nobody argues about it later.
Integrate & Build
Integration first, because that is where the surprises live. Then the AI layer on top. Working software in your repository from the first week, reviewable by your team at any point.
Evaluate & Harden
Scored against your historical cases. Escalation paths, failure handling, monitoring and cost controls built and tested. This is where the schedule is really spent, and it is the phase that most quotes leave out.
Launch & Hand Over
Staged rollout with humans in the loop, tightening as the numbers hold. Then the runbook, the training and a supported-operation period where your team drives and we back them up.
Before and After
Implementation works best when the target was chosen properly and someone owns it afterwards.
AI Readiness Assessment
Not sure what to build first? The $3k audit ranks the options and names the one worth starting with.
Start with the auditAI Training for Teams
A system nobody understands gets worked around. We train the people who will use and own what we build.
Team trainingFractional AI Officer
Senior oversight after handover, at a fraction of a full-time hire. Optional, and never a condition of the build.
Ongoing oversightFrequently Asked Questions
The questions that matter when you are buying a build rather than a slide deck.
Because a demo and a production system are different products, and the gap between them is where all the work is. A demo needs to succeed once, on a clean input, with someone knowledgeable driving. A production system has to handle the invoice that was photographed at an angle, the customer who asks three questions in one sentence, the API that times out at 4pm on the last day of the month, and the edge case nobody documented because everyone in the business just knows about it. The pilot gets to 80 per cent in three weeks, and the remaining 20 per cent — error handling, evaluation, integration, monitoring, the human escalation path — is 80 per cent of the actual engineering. Teams that have only ever built demos consistently underestimate it, then run out of budget and enthusiasm at the same moment.
Evaluation harnesses, built before the system goes anywhere near a customer. We construct a test set from your real historical cases — including the messy ones — and score every change against it, so we can tell you the accuracy on your data rather than a vendor benchmark. In production we log every interaction, sample outputs for human review, and alert on the things that actually matter: escalation rate climbing, confidence dropping, latency creeping, cost per transaction drifting. "It seemed fine when we tested it" is not a quality process, and it is the standard most AI deployments are held to.
This is the single most common hard constraint we hit in Australian mid-market businesses, and it is usually an industry vertical package whose vendor has no commercial interest in opening it up. There are several ways through: a supported integration or export you did not know existed, a database-level read where the licence permits it, a file-based exchange on a schedule, or robotic automation over the interface as a last resort — which we treat as a genuine last resort because it is brittle and breaks on UI updates. Occasionally the answer is that the system is the blocker and the honest recommendation is to change the system, not to build an elaborate workaround around it. We would rather tell you that in week one than bill you to build the workaround.
We hand over, and the engagement is structured to make that real. The code lives in your repository under your ownership from day one. The infrastructure runs in your accounts, on your billing, not resold through us. We write the runbook, we train your team on it, and we run a period of supported operation where your people drive and we sit behind them. Some clients then keep us on a light retainer for changes and oversight — see the fractional AI officer arrangement — and plenty do not. Both are fine. A build you cannot maintain without us is a liability we have sold you, not an asset.
A single well-scoped workflow — document extraction, an internal knowledge assistant, an enquiry triage agent — typically runs six to ten weeks from kick-off to production, including the evaluation and handover phases that most quotes quietly omit. Multi-workflow programmes or anything touching a system with a difficult integration surface run longer. We deliberately do not quote two-week AI builds, because a two-week build is a demo, and you would be paying us to create the exact pilot-purgatory problem you hired us to avoid.
Whichever ones fit the problem, and we are not aligned to any of them. We have no reseller agreements, no partner tiers and no commissions to protect, which means model selection is an engineering decision rather than a commercial one. In practice most systems end up using more than one model — a capable one for the hard reasoning step, a small cheap fast one for classification and routing — because that is usually the sensible cost and latency trade-off. We also design for substitution: the model sits behind an interface so you can swap it when something better or cheaper appears, which in this market it will, probably within months.
Got a Pilot That Never Shipped?
Bring it to the free consultation. We will give you a straight read on what it would take to finish it — or whether it is worth finishing at all.