Skip to main content

What are you looking for?

Explore our services and discover how we can help you achieve your goals

From AI proof of concept to production service

A validated pilot becomes a production service through one gate: fixed evaluation thresholds, confirmed data and access, a built integration, a named operations owner and rollback trigger, a recorded go or no-go decision, a watched rollout, and a close-out to ongoing operations. Skipping a step turns a demo into an unowned dependency.

Submit a project Assess the workload first

Reviewed by David Nguyen (CEO) · Updated 2 Oct 2026 · 10 min read

star

OutsourcingVN is operated by Netbase JSC, which builds and promotes AI pilots of this kind, so we have an interest in your decision. The gate below applies whoever runs it. It is written for CTOs, platform owners and product leads holding a pilot that works and asking whether it is ready to carry real traffic.

Contents

Which signal says promote, keep piloting, or stop?

Signal Promote to production Keep piloting Stop
Evaluation result Clears a fixed threshold on a held-out case set Close but not yet on the hardest cases Fails even on ordinary cases
Data and access Production sources and roles are confirmed and governed A source is still a one-off export The data the pilot needs cannot be reached safely
Integration A defined, reviewable connection to the systems it reads and writes Still a notebook or a manual copy-paste step No integration path exists at all
Operations ownership A named owner accepts the pager and the rollback trigger Nobody has agreed to own it yet No plausible owner exists in the organisation
Running-cost model Cost per request is measured and budgeted Cost is estimated but unmeasured Cost scales in a way nobody can afford to run

A pilot that is strong on evaluation but has no named operations owner is not ready; promoting on evaluation alone is the single most common way a demo becomes an unowned production dependency. The production AI inference assessment settles the serving route and acceptance cases this gate assumes are already decided; this guide is the promotion pipeline that follows once that workload boundary is fixed. Whether to run the pilot at all, before any of this, is covered in the emerging technology adoption guide.

How does a pilot become a production service?

  1. Freeze evaluation thresholds

    Fix the acceptance thresholds the pilot must clear before promotion; loosening a threshold during the review meeting defeats the point of setting one in advance. The AI workflow evaluation and testing guide covers building the case set this step scores against.

  2. Confirm data and access

    Name the production data sources, the access roles allowed to call the service, and what happens when a dependency is unavailable.

  3. Build the integration

    Connect the service to the systems it must read from and write to, behind the same version control and review process as any other production change.

  4. Name the operations owner and the rollback trigger

    Agree in advance which signal after launch means roll back now, so nobody debates it mid-incident, and name who owns the pager.

  5. Decide and record the gate

    The product owner, the operations owner and the risk or compliance owner sign off together: go, go with recorded exceptions, or no-go with a new date.

  6. Deploy and watch

    Release behind a flag or to a small share of traffic first, watching the agreed signals for an agreed monitoring period before a full rollout.

  7. Close the promotion

    Record what the gate found, confirm the running-cost model matched reality, and hand the validated pilot to its ongoing operations owner.

Google Cloud's own MLOps guidance describes the same shape from the infrastructure side: a newly promoted model undergoes online validation in a canary deployment or an A/B test before it serves full traffic, and teams keep a pointer to the previous model version so they can roll back if the new one underperforms.

Diagram of a pilot-to-production pipeline: freezing evaluation thresholds, confirming data and access and building the integration lead to a go or no-go decision, then a watched deployment with a rollback edge back (opens the full-size diagram in a new tab)
Diagram of a pilot-to-production pipeline

A worked scenario: a document-extraction pilot reaches its gate

A logistics company's data team built a pilot that extracts fields from delivery paperwork, tested on a folder of saved documents, and it clears every example by hand. The operations manager asks if it is ready to run every night.

Walking the gate finds three gaps before a yes: the evaluation case set has no deliberately malformed or ambiguous documents, so the threshold is frozen only after adding them; the pilot currently reads from a shared drive rather than the overnight document queue, so the integration step connects it to that queue with retry and idempotency; and nobody has agreed who is paged if extraction accuracy drops after a supplier changes its invoice layout. Once the operations manager accepts that pager and a rollback trigger is agreed — reverting to manual entry if the error rate crosses a set level — the service deploys behind a flag against a small nightly batch, watched for two weeks, before it takes the full queue. Cloud Platform and DevOps Engineering is the service that builds the queue, observability and cost controls this kind of promotion depends on, often on AWS; the reliability and incident readiness guide covers what happens once an incident does occur, and the edge AI deployment assessment covers the separate case where the model must run on a device rather than in the cloud.

What has Netbase delivered near this area?

Netbase has delivered anonymised client AI projects including retrieval-based knowledge assistants, document AI and MLOps pipelines. A published record that moved through milestones comparable to this gate is the WhatsApp AI chatbot with CRM integration: the chatbot was delivered in milestones from design and prototype to documentation, knowledge transfer and 30 days of support over four to eight weeks, the same design-prototype-handover shape this guide's gate formalises for a production release. Longer-running operations discipline shows in the multi-tenant cloud ERP SaaS platform, where Netbase has worked as offshore development and managing partner since 2020 on a stack that includes React and Next.js, Laravel and Strapi, and PostgreSQL, MySQL, MariaDB and MongoDB on AWS.

Netbase delivers remote-first from Hanoi in Agile increments with weekly reviews, using AI-assisted engineering under human review, and its delivery lifecycle runs discovery and strategic alignment, team assembly and architecture planning, agile execution with outcome-based milestones, modular components, training and rollout, and ongoing support; a pilot's promotion gate is that lifecycle's training-and-rollout and ongoing-support steps made explicit for an AI service. Security practices include secure code review and version control, role-based access control, multi-factor authentication for admin dashboards, contributors under NDA, and NDAs and data processing agreements on request, applied to the data and access step above. Netbase applies the ISO/IEC 42001 AI management system framework to its own AI delivery practice, which governs how a promoted model is evaluated and rolled back; many pilots that reach this gate started as an engagement scoped under AI Workflow Automation.

Which questions should you ask a supplier?

  • What threshold must the pilot clear, and who set it?

    A number on a held-out case set, agreed before the review, not during it

  • Who owns the service after go-live?

    A named person who accepts the pager, not "the team"

  • What is the rollback trigger?

    A specific signal agreed in advance, with the previous version kept ready

  • How is the running cost measured?

    Cost per request or per batch, watched against a budget from day one

  • What happens on a no-go?

    A recorded reason and a new date, not a quiet retry next sprint

  • How is the rollout staged?

    A flag or a small traffic share first, watched before a full release

What goes wrong when a pilot is promoted without this gate?

  • Evaluation on the easy cases only. Signal: the pilot fails on its first real edge case in production. Owner: the product owner, who adds malformed and ambiguous cases before the threshold is frozen.
  • No named operations owner. Signal: an alert fires and nobody responds. Owner: the operations owner, agreed and accepted before go-live, not assumed after.
  • No rollback trigger agreed in advance. Signal: a mid-incident argument about whether to roll back. Owner: the operations owner, who has the trigger and the previous version ready beforehand.
  • Running cost discovered on the bill. Signal: a usage spike surprises the finance owner. Owner: the platform owner, who wires a budget alert before the rollout, not after.
  • A full rollout with no staged step. Signal: every user hits the new pilot at once. Owner: the release owner, who ships behind a flag or to a small share of traffic first.

How this guide is sourced

Statements about Netbase come from attested company facts listed under Sources. External guidance on staged rollout, evaluation and rollback comes from Google Cloud's own MLOps architecture documentation, dated below. The document-extraction scenario is illustrative and describes no client. This page does not estimate evaluation thresholds, running cost or rollout risk for your specific pilot. See the methodology for how claims are reviewed.

Plan the next step for your project

Common questions

Rarely. A prototype usually runs on clean inputs at low concurrency and answers a narrower question than production traffic asks. Treat promotion as a new release with its own evaluation, rollback path and named operator.

The product owner who accepts the behaviour, the operations owner who accepts running it, and a risk or compliance owner where the data or decision warrants one. A meeting without the operations owner is not a real gate.

Record a no-go and a new date, fix the specific gap the evaluation found, and rerun the same case set rather than a looser one. Lowering the threshold to pass is not a fix.

Long enough to see the failure modes that would show up in real use, often one to a few weeks depending on how often the workload runs; a nightly batch job needs more cycles than a service that handles thousands of requests an hour.

The gate scales with consequence. A low-stakes internal tool can run a lighter version of the same steps; a service touching customer data or money should not skip any of them.

Bring your pilot and its evaluation results

Bring the pilot, its evaluation results so far, and the gaps you already know about in data, access or ownership. Submit a project with that material, or read the production AI inference assessment first if the serving route itself is still open. OutsourcingVN is Netbase's own outsourcing-services platform.

Cloud platform and DevOps engineering: one platform for apps and AI workloads Cloud platform and DevOps engineering: one platform for apps and AI workloads

Infrastructure as code, CI/CD, inference hosting, observability and cost controls for apps and AI workloads.

Learn More
line

Tell us what you want to build or automate.

Submit a project