RAG vs fine-tuning is often presented as a choice between two ways of making an AI model understand your business.
That is only partly correct.
Retrieval-augmented generation, fine-tuning, live system integrations and the choice between a local or frontier model solve different problems.
If a company wants an AI assistant to know its policies, products, documentation and internal processes, training all of that information into the model is usually not the first thing we would recommend.
In many cases, the better starting point is to keep company knowledge outside the model and retrieve the relevant information when it is needed.
Fine-tuning then becomes useful when the problem is not what the model knows, but how consistently it performs a particular task.
And if the AI needs today’s customer balance, current stock level or the ability to update a CRM, neither RAG nor fine-tuning is enough by itself. The model needs controlled access to the relevant business system.
This distinction becomes particularly important as companies move towards locally deployed AI. A business does not necessarily need to build or train its own foundation model to create AI that understands how the organisation works.
RAG vs fine-tuning: the simplest distinction
The easiest way to separate the two is this:
RAG gives the model information at the time it needs it.
Fine-tuning changes how the model itself behaves.
| Approach | What it changes | Useful for |
|---|---|---|
| RAG | The information supplied to the model during a request | Company documents, policies, product information, knowledge bases |
| Fine-tuning | The model’s learned behaviour | Consistent formats, classification, specialist task behaviour, instruction following |
| Tools / integrations | What live data and actions the model can access | CRM records, stock, bookings, account information, business actions |
| Frontier model | The underlying level of general model capability | Difficult reasoning, complex coding and tasks where stronger general intelligence materially improves the result |
The approaches are not mutually exclusive.
A local model can use RAG.
A fine-tuned model can use RAG.
A frontier API can use RAG.
A locally hosted model can call company tools.
And a system can combine all of them while routing particularly difficult requests to a stronger model.
That is why the real architectural question is not simply “RAG or fine-tuning?”
It is:
What exactly are we trying to make the AI know, do or access?
If the AI needs company knowledge, start by thinking about RAG
Imagine a company wants an internal assistant that can answer questions about:
- HR policies;
- technical documentation;
- product specifications;
- sales procedures;
- contracts;
- customer-support guidance;
- internal processes;
- or operational manuals.
That information changes.
A policy gets updated. A new product launches. A procedure is replaced. A document is withdrawn.
Training those facts directly into model weights creates an awkward maintenance problem: when the information changes, the model may need to be fine-tuned again.
Retrieval-augmented generation takes a different approach.
The documents remain in a knowledge system. When somebody asks a question, the application searches for relevant information and supplies that material to the model as context before it produces its answer.
The model does not need to permanently memorise every company document.
It needs access to the right information at the right time.
AWS’s current guidance comparing RAG and fine-tuning recommends starting with RAG when building question-answering systems around custom documents. AWS specifically notes that RAG can incorporate updated documents quickly without retraining the underlying model.
For company knowledge, that is a major advantage.
RAG means your knowledge can change without retraining the model
Consider an employee handbook.
If the holiday policy changes on Monday, an internal AI assistant should ideally be able to use the new policy on Monday.
With RAG, the underlying document can be updated and re-indexed.
The model itself does not necessarily need to change.
The same principle applies to:
- product documentation;
- price lists;
- technical procedures;
- compliance guidance;
- internal FAQs;
- service information;
- and frequently changing operational knowledge.
This separation between model capability and company knowledge is useful.
The model provides the language and reasoning capability.
The business controls the information supplied to it.
That also makes it easier to understand where an answer came from. A well-designed RAG system can retain references to the source material used to construct an answer rather than expecting employees to trust information apparently recalled from a model’s weights.
RAG is not the same as simply giving an AI access to every company file
A useful RAG system still has to be engineered properly.
Uploading thousands of documents into a database and pointing an AI model towards them does not guarantee useful answers.
The system needs to decide:
- which documents should be indexed;
- how they should be divided into retrievable sections;
- how information is searched;
- how many results are supplied to the model;
- which sources should take priority;
- how old information is removed;
- and which employees are allowed to retrieve which documents.
That last point is especially important.
If an employee should not be able to read a confidential HR document directly, an AI assistant should not be able to retrieve it on their behalf either.
RAG needs to respect the same permissions as the systems around it.
This connects directly with the architecture discussed in Private AI: Should Your Company’s AI Run Inside Its Own Infrastructure?. If both the model and company knowledge store are operated inside controlled infrastructure, sensitive internal information does not necessarily need to be sent to a public AI service to answer an employee’s question.
What fine-tuning actually solves
Fine-tuning is useful for a different class of problem.
Instead of continually giving the model information to reference, the developer provides examples that teach it to behave more consistently on a particular task.
OpenAI’s supervised fine-tuning guidance describes use cases including classification, producing consistent formats and correcting instruction-following problems.
Imagine a company needs an AI system to read incoming support messages and assign exactly one of 40 internal categories.
The problem may not be missing knowledge.
The model may simply need to become much more consistent at that particular classification task.
Or perhaps an AI needs to turn unstructured inspection notes into an exact internal JSON structure every time.
Again, that is principally a behaviour problem.
Fine-tuning can make sense when repeated prompting is not producing consistent enough results and the company has enough high-quality examples of what correct behaviour looks like.
Fine-tuning is usually a poor substitute for a changing knowledge base
This is where companies can spend money solving the wrong problem.
Suppose the objective is:
“We want the AI to know everything in our company documentation.”
Fine-tuning might sound like the obvious solution because the word “training” suggests teaching the model the information.
But if those documents change regularly, training the information into model behaviour can make updates considerably harder.
A RAG system can replace or update a source document.
A fine-tuned model generally needs another tuning process if the information embedded through training needs to change.
AWS also notes another important distinction: fine-tuned models do not naturally provide references back to the source documents they learned from, while RAG systems can retain the source used for a particular answer.
For internal knowledge systems, that can matter a great deal.
If someone asks an AI assistant about a contractual process or company policy, being able to inspect the underlying document is often more useful than receiving an answer with no visible provenance.
Fine-tuning a local model can still be extremely useful
None of this means companies should avoid fine-tuning.
It means fine-tuning should solve the right problem.
Open-weight models have made that particularly interesting because fine-tuning no longer has to mean sending training examples to a closed provider and then depending permanently on its hosted fine-tuning service.
OpenAI states that its gpt-oss open-weight models can be fine-tuned using open tooling and infrastructure controlled by the organisation.
That allows a company to take a capable general model and adapt it for a narrower internal role while retaining control over where inference takes place.
For example, a company could fine-tune a local model to:
- classify its own types of enquiries;
- follow a specialist workflow;
- produce a required internal response format;
- use organisation-specific terminology consistently;
- or perform a repeated narrow task more reliably.
This is a much more realistic interpretation of “custom AI” than trying to create a foundation model from scratch.
It also continues the economic argument from Self-Hosted AI: Is It Cheaper Than OpenAI or Anthropic?. If a specialist model handles a high-volume internal task reliably, the organisation can operate that capability on infrastructure it controls rather than paying a frontier provider for every inference.
There is an important current change in hosted fine-tuning
The distinction between owning a model and relying on a provider’s customisation service is also becoming more visible.
OpenAI’s current documentation says its hosted supervised fine-tuning platform is being wound down and is no longer available to new users, while existing fine-tuned models remain available during the transition according to their base-model lifecycle.
At the same time, OpenAI continues to support fine-tuning of its open-weight gpt-oss models using external tooling.
This should not be interpreted as evidence that hosted fine-tuning is disappearing across the entire AI industry.
It does demonstrate a wider architectural point, though:
If an organisation’s custom AI capability depends entirely on a provider-specific product, that capability also depends on the provider continuing to offer it.
A self-hosted open-weight model gives the business more control over that lifecycle.
What if the information is live?
RAG is useful for documents and knowledge.
But not all company information belongs in a document store.
Suppose an employee asks:
“Has customer 4821 paid their latest invoice?”
Or:
“How many units of this product are currently in stock?”
Or:
“Move this support ticket to the escalation queue.”
Those are not really RAG questions.
The answer belongs in a live business system.
The AI needs controlled access to the accounting platform, stock database, CRM, helpdesk or another source of current information.
This is where tools and function calling become important.
OpenAI’s current function-calling documentation describes this architecture as giving a model access to external data and functions through application-controlled tools.
The same concept applies whether the underlying model is hosted externally or running locally.
The model decides it needs information or an action.
The application performs the authorised operation.
The result is then returned to the model.
In other words:
Do not train today’s stock level into a model.
Do not rely on a RAG index of yesterday’s CRM export either.
Connect the AI to the actual system through a controlled interface.
A useful company AI therefore has several kinds of “knowledge”
It helps to separate them.
| What the AI needs | Likely approach |
|---|---|
| General language and reasoning ability | Base model |
| Company documents and internal knowledge | RAG |
| Consistent specialist behaviour | Fine-tuning where justified |
| Current account, CRM or operational information | Tools / system integrations |
| Ability to perform business actions | Controlled tools with permissions |
| Exceptionally difficult reasoning | Frontier model escalation |
That is a much more useful way to think about custom business AI than attempting to place everything inside a single model.
Where does a frontier API fit into RAG vs fine-tuning?
This is another point that is frequently confused.
A frontier API is not really an alternative to RAG or fine-tuning.
It is one possible place from which the underlying model capability comes.
You can build RAG around a frontier API.
You can build RAG around a local open-weight model.
You can fine-tune an appropriate model and then add RAG.
You can give either a local or hosted model access to tools.
The choice of model and the choice of knowledge architecture are separate decisions.
This is exactly why we argued in AI Model Routing: Why One Model for Your Company Isn’t Enough that businesses should stop treating one model as the automatic destination for every request.
A company may use a local model with RAG for most internal knowledge questions and escalate unusually difficult requests to a stronger external model.
The company knowledge system can remain largely unchanged.
Only the model processing a particular request changes.
The architecture we would normally investigate first
For many businesses building an internal AI system today, we would investigate the components in roughly this order:
- Start with a capable model appropriate to the workload. Do not assume it needs to be the most powerful model available.
- Add RAG for company knowledge. Keep changing policies, documentation and reference material outside the model.
- Add controlled tools for live systems. Let the application fetch current information and perform authorised actions.
- Evaluate the real workload. Find where the model is failing rather than assuming custom training is required.
- Fine-tune when the failures are behavioural and repeatable. Use good examples to improve a narrow task.
- Escalate genuinely difficult requests when needed. A frontier API can remain available without becoming the default for every inference.
This approach avoids unnecessary training while preserving the option to specialise the model when there is evidence that specialisation will help.
Why a local model plus RAG can be such a strong business combination
For the series of articles we have published around local AI, this is where the pieces come together.
A smaller local model may be capable enough for a large proportion of company work.
RAG gives that model access to the organisation’s changing internal knowledge without needing to encode that knowledge permanently into the model.
Tools connect it to current business systems.
Fine-tuning can improve repeated specialist behaviour.
And model routing can send exceptional requests elsewhere when they genuinely need more capability.
That produces an architecture that can look something like this:
Employee request
|
v
Permissions and request classification
|
+---- Needs company knowledge? ----> Retrieve approved internal information
|
+---- Needs live data/action? -----> Call authorised business tool
|
v
Local company model
|
+---- Result passes requirements ---> Return answer
|
+---- Requires stronger reasoning --> Escalate to approved frontier model
This is more involved than simply connecting every internal application directly to one AI API.
But it also gives the company substantially more control over cost, data, behaviour and provider dependency.
Do not fine-tune before you can explain what is wrong
Fine-tuning sounds attractive because it makes an AI system feel genuinely bespoke.
That is not a good enough reason to do it.
Before training anything, the development team should be able to describe the problem clearly.
For example:
Bad reason: “We want the AI to understand our company better.”
Better diagnosis: “The model has access to the correct policy but repeatedly fails to apply these three classification rules.”
The second problem can be measured.
Examples can be collected.
A fine-tuning dataset can be created.
The tuned model can then be compared against the original model to determine whether anything actually improved.
If the real problem was that the AI could not find the latest policy, fine-tuning would have been solving the wrong problem.
RAG needs evaluation too
The same discipline should apply to retrieval.
A poor answer does not automatically mean the model is too weak.
Perhaps the retrieval system supplied the wrong document.
Perhaps an outdated version ranked above the current one.
Perhaps the relevant section was split badly when the document was indexed.
Perhaps the employee did not have permission to retrieve the necessary source.
Perhaps the answer required live system data rather than documentation.
Good company AI therefore needs evaluation across the complete system, not just the language model.
That is another reason we are cautious about treating model benchmarks as the main measure of business AI quality.
The strongest model in the world cannot answer from a company policy it has never been given.
So should your business use RAG, fine-tuning or a frontier model?
For company knowledge that changes, start with RAG.
For repeatable behaviour that remains inconsistent after good prompting and system design, consider fine-tuning.
For live company information and actions, connect the model to controlled tools and business systems.
For unusually difficult work where a local model genuinely falls short, retain access to a frontier model.
And if more of the workload can be handled by a locally deployed open-weight model, there is no requirement for every one of those components to sit inside somebody else’s AI platform.
Our view at Nort Labs is that this is a more useful definition of custom AI.
It is not necessarily a model trained from scratch on everything the company knows.
It is an AI system designed around the company’s knowledge, permissions, workflows, systems and actual level of required intelligence.
Sometimes that means RAG.
Sometimes it means fine-tuning.
Often it means integrations.
And increasingly, much of it can run inside infrastructure the organisation controls.
Nort Labs can design local and hybrid AI systems that combine company knowledge, RAG, fine-tuned models, live business integrations and frontier-model escalation around the requirements of the organisation rather than forcing every workload through a single external AI service.