OutsourcingVN is operated by Netbase JSC, which builds and evaluates both retrieval and fine-tuning pipelines and would like to build yours, so read this as a supplier's method. The questions work with whoever builds it.
Contents
- What problem is each lever actually solving?
- Which lever fits your problem?
- Which approach fits your problem?
- Worked scenario: a support team's two separate requests
- What has Netbase delivered near this area?
- Which questions should you ask a supplier?
- What goes wrong when teams choose the wrong lever?
- How this guide is sourced
- Common questions
- Bring your workflow and your data
What problem is each lever actually solving?
Retrieval-augmented generation (RAG) leaves a model's weights untouched and instead fetches relevant passages at answer time, so it answers from text it was never trained on. Fine-tuning adjusts the model's own weights on a set of labelled examples, so the change persists without a retrieval step, but only for what the training data captured. The two solve different problems: retrieval keeps facts current and attributable to a source; fine-tuning changes how a model behaves — its tone, output format, task-following or domain vocabulary — not what it factually knows.
This distinction matters because the two are often confused with the underlying technology choice. The large language models page covers where either pattern fits inside a wider product, and the RAG context engineering guide scopes a retrieval project once source authority and access rules are the open question, not whether retrieval is the right lever at all. A related but separate decision is ranking or matching records rather than grounding an answer, covered in the embeddings and semantic search guide.
Which lever fits your problem?
| Criterion | Retrieval (RAG) | Fine-tuning | Both | Neither |
|---|---|---|---|---|
| Knowledge freshness | Facts change often, or every answer must cite a source | Facts are stable and already inside the base model | Facts change; style or format must also change | No outside knowledge is needed at all |
| Behaviour change needed | None; the base model already writes acceptably | Tone, format, task-following or domain vocabulary must change | Both knowledge and behaviour must change | A single well-written prompt already works |
| Data volume available | A source register of documents; no labelled examples required | Hundreds to thousands of labelled input-output examples | Both a source register and labelled examples | Neither a usable source nor enough examples |
| Evaluation effort | A judgement set of real questions with expected citations | A held-out test set scored against the target behaviour | Both evaluation sets, run and tracked separately | A handful of example prompts |
| Running cost | Retrieval latency and an index to operate | A training run, repeated as data changes, then ordinary inference | Both an index and training runs to maintain | Base model cost only |
Most production systems that need both choose retrieval for facts and a lighter fine-tuning or prompting pass for behaviour, because fine-tuned knowledge goes stale and cannot be cited, while prompting alone eventually runs out of room for complex formatting rules. OpenAI's own guidance frames fine-tuning as the later step once prompt engineering alone no longer gets acceptable results, and best suited to tasks such as classification, strict formatting or correcting instruction-following failures rather than teaching a model new facts.
Which approach fits your problem?
-
Do the facts change often, or must every answer be traced to a source?
If yes, retrieval (RAG) is the lever: fine-tuned knowledge goes stale and cannot be cited, so facts belong in a retrieval index, not training data.
-
Must the model's behaviour change — its format, tone, task-following or domain vocabulary — while the facts stay stable?
If yes, fine-tuning alone is often enough: it changes how the model writes, not what it needs to know.
-
Do both the facts and the behaviour need to change, and do you have hundreds of labelled examples to show the behaviour you want?
If yes, combine retrieval for the facts with fine-tuning for the behaviour; each pipeline is evaluated on its own test set.
-
Would a single well-written prompt already pass your evaluation set?
If yes, build neither a fine-tuning pipeline nor a bespoke retrieval architecture yet: ship the prompt, measure it, and add either lever only when a measured gap remains.
Whichever branch you land on, the AI workflow evaluation and testing guide covers how to build and maintain the judgement or test set each lever depends on; a lever adopted without one cannot be shown to help.
(opens the full-size diagram in a new tab)
Worked scenario: a support team's two separate requests
A software company's support team asks for two things in the same project brief: answers that quote the current help-centre articles, and replies that always follow the company's three-part tone and ticket-closing format.
Walking the questions above splits the request. The help-centre answers depend on articles that change weekly, so they need retrieval: a source register, an index and a judgement set of real tickets, the pattern a single AI knowledge assistant can be scoped to build. The tone and format requirement does not depend on facts changing, so it is tested first as a prompt instruction; only once a held-out set of real replies shows the prompt still drifts from the required format does a small fine-tuning run get scoped, on the hundreds of past replies the team already has. The two pipelines are built, evaluated and released separately, even though both land in the same chat window.
What has Netbase delivered near this area?
Netbase has delivered anonymised client AI projects including retrieval-based knowledge assistants, document AI and MLOps pipelines; retrieval sits behind each of those, applying the same source-register and evaluation discipline this guide describes. A published record in the generative direction is the WhatsApp AI chatbot with CRM integration: for a client (not named), Netbase built a WhatsApp Business AI chatbot with intent and conversation-flow handling, an LLM API (GPT-4 in the stack) and CRM synchronisation for lead capture, customer data and workflow automation; that chatbot used a hosted model with prompting, not a fine-tuned one.
Netbase works with commercial and open-source AI models chosen per project (model-agnostic); no vendor partnership is implied, whether the model is retrieved against or selected as the base for fine-tuning. The model selection guide covers choosing that base model for one workflow. Netbase applies the ISO/IEC 42001 AI management system framework to its own AI delivery practice, which governs how a project evaluates and can roll back either a retrieval change or a fine-tuned model version. Where a second or third workflow needs the same governed source register, index and evaluation store rather than a one-off build, data and AI platform engineering is the shared foundation; whether the underlying data is even ready for either lever is the separate question the data platform readiness assessment answers first.
Which questions should you ask a supplier?
-
Why this lever and not the other?
A mapping to the criteria above, not a default choice
-
How many labelled examples does the fine-tuning case need?
A number tied to the task, with a plan if you have fewer
-
How is the retrieval index kept current?
A refresh trigger tied to source changes, not a calendar guess
-
How is our data protected during fine-tuning?
Security practices including secure code review and version control, role-based access control, multi-factor authentication for admin dashboards, contributors under NDA, and NDAs and data processing agreements on request
-
What compliance rules apply to the training or source data?
GDPR alignment for data privacy in Europe, HIPAA-aligned methodologies for healthcare data handling, and CCPA compliance for clients with U.S. customer bases, as practices rather than a guarantee of your own legal compliance
-
How do you roll back a fine-tuned model version?
A pinned previous version and a re-test before any swap
-
What happens if the chosen lever does not clear evaluation?
A named fallback, not a shipped feature regardless
What goes wrong when teams choose the wrong lever?
- Fine-tuning to fix stale facts. Signal: the model keeps citing outdated figures after a retrain. Owner: the data owner, who moves the facts into a retrieval index instead of another training run.
- Retrieval used to fix tone or format. Signal: retrieved passages are correct but the reply still ignores the house style. Owner: the product owner, who scopes a prompt or fine-tuning pass for behaviour.
- Fine-tuning on too few examples. Signal: the model overfits a handful of cases and performs worse on new ones. Owner: the engineer, who falls back to prompting against a retrieval index until enough examples exist.
- No held-out test set for either lever. Signal: nobody can say whether the new version is actually better. Owner: the product owner, who insists on a fixed evaluation set before any release.
- Treating the two levers as mutually exclusive forever. Signal: a workflow that clearly needs both ships with only one. Owner: the workflow owner, who scopes and evaluates each pipeline separately rather than forcing one lever to do both jobs.
How this guide is sourced
Statements about Netbase come from attested company facts listed under Sources. External guidance on when fine-tuning fits and its data requirements comes from OpenAI's and AWS Bedrock's own customization documentation, both dated below. The support-team scenario is illustrative and describes no client. This page does not estimate training cost, data volume or evaluation results for your specific case. See the methodology for how claims are reviewed.
Plan the next step for your project
Common questions
Not reliably. Facts learned during fine-tuning are baked into the weights, go stale as the world changes, and cannot be cited back to a source. For facts, retrieval stays necessary even after a fine-tuning project.
Yes, when the retrieved passages are right but the reply's tone, format or task-following still misses the mark. The two are evaluated and released as separate pipelines even when a single feature uses both.
Supplier documentation describes the workflow rather than a fixed number; in practice it depends on the task, and a few dozen examples rarely clears a held-out evaluation set the way hundreds to thousands usually do.
Usually to start, because no training run or retrained-model pipeline is needed. Fine-tuning can lower per-request cost later by shortening prompts, once a measured gap justifies the extra pipeline.
Yes. The two are independent choices about how a specific feature uses the model; a model used for retrieval-grounded answers elsewhere can still be the base for a separate fine-tuning project.
Bring your workflow and your data
Bring the workflow, a sample of the documents or examples involved, and which problem — stale facts or wrong behaviour — is actually biting. Submit a project with that material, or read the RAG context engineering guide and data platform readiness assessment first if the underlying sources are still unclear. OutsourcingVN is Netbase's own outsourcing-services platform.
Related services and solutions
Data and AI platform engineering: one foundation your workflows share
Pipelines, retrieval indexes and model operations your AI and analytics workflows can share.
Learn More