OutsourcingVN is operated by Netbase JSC, which scopes and builds AI features and would like to build this one, so read the guide as a supplier's method. The questions work with any team.
Contents
- What problem do embeddings actually solve?
- Which feature needs embeddings, and which does not?
- How do you judge relevance before launch?
- How should the index be built and refreshed?
- Worked scenario: semantic search for a supplier directory
- What has Netbase delivered near this area?
- Which questions should you ask a supplier?
- What goes wrong with semantic search?
- How this guide is sourced
- Common questions
- Bring your queries and the feature you want to improve
What problem do embeddings actually solve?
An embedding model turns a piece of text, an image or a record into a list of numbers so that items with similar meaning sit close together. Search then becomes "find the nearest items to this query", which catches synonyms, paraphrases and misspellings that keyword search misses. The same geometry supports several product features, and each needs a different definition of "close".
This guide is about embeddings as a product feature: the search box, the matching engine, the duplicate detector. When the goal is to ground a chatbot's answers in approved documents, the RAG context engineering guide covers source authority, citations and answer evaluation instead. The emerging technology adoption guide frames where both sit among other AI decisions.
Which feature needs embeddings, and which does not?
| Feature | Embeddings help when | Keyword or rules are enough when | What "relevant" means |
|---|---|---|---|
| Site or catalogue search | Users describe needs in their own words ("quiet laptop for travel") | Users search by exact codes, names or SKUs | The item a reasonable user would open first |
| Matching (jobs, suppliers, listings) | Descriptions are free text on both sides | Matching is on structured fields such as location and category | A match the operator would accept without editing |
| Recommendation ("similar items") | Similarity of content matters more than purchase history | Behaviour data is rich and content is thin | An item a user plausibly considers next |
| Deduplication | Duplicates are reworded, not copied | Duplicates share an identifier or exact text | Two records an editor would merge |
| Ticket or message routing | Categories are fuzzy and messages are short | Senders choose a category reliably | The queue a supervisor would have chosen |
Most production systems end up hybrid: keyword search for exact identifiers, embeddings for meaning, and filters for hard constraints such as location, stock or permission. Plan for a hybrid from the start rather than replacing keyword search wholesale.
How do you judge relevance before launch?
-
Collect real queries
Take a few hundred from search logs, support tickets or sales conversations. Invented queries flatter the system.
-
Write the judgement rule
One sentence per feature, taken from the table above, agreed by the product owner.
-
Label a judgement set
For each query, have two people mark the top results as relevant, partly relevant or not relevant. Where they disagree, refine the rule rather than averaging.
-
Measure the baseline
Score your current keyword search on the same set. Standard ranking measures such as precision at the first few results and whether the best item appears near the top are described in the Stanford Introduction to Information Retrieval text.
-
Compare candidates on the same set
Try two or three embedding models and a hybrid. Keep the judgement set fixed so the comparison is fair.
-
Set a release threshold
Decide in advance how much better the new ranking must be, and which query groups must not get worse.
-
Keep the set alive
Add new failing queries every month so the evaluation tracks what users actually type.
How should the index be built and refreshed?
An index is only as good as what goes into it. Decide these points in the specification, not during testing:
- Unit of meaning. A whole product record, a listing description, or a paragraph of a long document. Too large and matches blur; too small and results lose context.
- Fields to embed and fields to filter. Embed descriptive text; keep stock status, locations, dates and permissions as structured filters applied before or after the vector search.
- Refresh trigger. Re-embed a record when its descriptive text changes, and remove it from the index when it is withdrawn. A stale index returns items that no longer exist.
- Model change plan. Vectors from different models are not comparable. Changing the embedding model means re-embedding the whole collection, so budget the time and keep the old index until the new one passes the judgement set.
- Access control. If results depend on who is searching, the filter must run on every query. The OWASP Top 10 for LLM Applications 2025 lists vector and embedding weaknesses, including leakage across permission boundaries and poisoned entries, as a distinct risk.
Latency, hosting and cost of serving the model belong to the production AI inference assessment; plan both together when the search box sits on a busy page.
Worked scenario: semantic search for a supplier directory
A regional trade association runs a supplier directory with 8,000 listings. Members complain that searching "food-safe packaging" misses suppliers who describe themselves as "FDA-grade containers" or "compostable takeaway boxes".
- Week 1. The team exports 400 real queries from the search log and writes the rule: a result is relevant if a buyer would shortlist the supplier for that need.
- Week 2. Two association staff label the top ten results for each query under the current keyword search. The best supplier appears in the top three for fewer than half of queries.
- Week 3. Listing descriptions are embedded per listing; location, category and membership status stay as filters. A hybrid ranking combines keyword and vector scores.
- Week 4. On the same 400 queries, the hybrid places the best supplier in the top three far more often, but queries that name a company get slightly worse. The team routes exact-name queries to keyword search and ships behind a feature flag to a tenth of visitors.
The figures in this scenario are illustrative, not a Netbase result. The pattern matters: a fixed judgement set, a hybrid rather than a replacement, and a narrow first release. The same trust questions apply to any listing product, which the business directory platforms page covers from the operator's side.
What has Netbase delivered near this area?
Netbase has delivered anonymised client AI projects including retrieval-based knowledge assistants, document AI and MLOps pipelines. Retrieval-based assistants rest on the same embedding and indexing decisions described above, applied to answering questions rather than ranking products.
A published record in the conversational direction is the WhatsApp AI chatbot with CRM integration: for a client (not named), Netbase built a WhatsApp Business AI chatbot with intent and conversation-flow handling, an LLM API (GPT-4 in the stack) and CRM synchronisation for lead capture, customer data and workflow automation. Intent handling is a close cousin of routing: short messages mapped to the right path.
Netbase works with commercial and open-source AI models chosen per project (model-agnostic); no vendor partnership is implied. For embeddings this means the model is chosen on your judgement set and can be replaced later, with a planned re-index.
Which questions should you ask a supplier?
- What relevance rule will you test against, and who wrote it? A good answer names your product owner, not the supplier.
- How large is the judgement set, and where do the queries come from? Real logs beat invented examples.
- What stays keyword or filter based? A supplier who replaces everything with vectors has not looked at exact-match queries.
- How is the index refreshed, and what happens when a record is withdrawn? Expect a trigger per record, not a nightly full rebuild by default.
- What happens if we change the embedding model? The answer should include a full re-embed and a side-by-side test.
- How are permissions enforced on results? Filters must run on every query, including for administrators testing the system.
- What do you monitor after launch? Zero-result queries, click position and a monthly re-score of the judgement set.
If the answers are thin, an AI workflow blueprint can fix the judgement rule, data and success measure before a build is scoped.
What goes wrong with semantic search?
- Launching without a baseline. Signal: nobody can say whether the new search is better. Owner: the product owner, who insists on scoring keyword search first.
- Losing exact matches. Signal: users typing a product code get "similar" items. Owner: the engineer, who routes identifier-like queries to keyword search.
- Stale vectors. Signal: withdrawn listings keep appearing. Owner: whoever owns the content pipeline, who wires re-embedding to record changes.
- Silent model drift. Signal: a provider updates a hosted model and rankings shift. Owner: the engineer, who pins model versions and re-runs the judgement set before any change.
- Leaking across permissions. Signal: a user sees a result from another tenant. Owner: the security reviewer, who tests filters with hostile queries; the AI assurance and red-teaming guide describes that testing.
- Treating similarity as truth. Signal: deduplication merges two genuinely different records. Owner: the operator, who keeps a human review step for merges above a set impact.
How this guide is sourced
The method combines a standard information-retrieval text for relevance measurement, the OWASP 2025 list for vector and embedding risks, and Netbase records approved in the OutsourcingVN claim register: anonymised AI delivery, the WhatsApp chatbot and the model-agnostic approach. The worked scenario is illustrative, and its numbers are not measured results.
Plan the next step for your project
Common questions
Not always. Small collections can use a vector extension in the database you already run; a dedicated vector store earns its place when the collection or query volume grows, or when you need specialised filtering. Decide after the judgement set shows the feature is worth building.
Enough to cover your main query groups with a few dozen each. A few hundred real queries is a practical start for one feature, and the set should grow as new failures appear.
Multilingual embedding models place similar meanings close together across languages, so a query in one language can find a listing written in another. Test it with real bilingual queries, because quality varies by language pair.
Re-embed a record whenever its descriptive text changes, and re-embed everything when you change model. A scheduled full rebuild is a fallback, not the main mechanism.
No. Semantic search returns ranked items for a person to choose; a retrieval-augmented assistant writes an answer from retrieved passages and needs citation and refusal rules on top.
Bring your queries and the feature you want to improve
Start with a sample of real queries, the records they should find and the feature you want to change. OutsourcingVN is Netbase's own outsourcing-services platform. Submit a project with that material, and a person will reply with whether a bounded assessment or a build is the right next step.
Related services and solutions
AI workflow blueprint: decide where automation belongs before you build
A paid two-to-four-week project that ends in a scope you can approve, change or stop.
Learn More