Why This Matters for Business
Off-the-shelf AI models are generalists. They can draft emails and summarize documents, but they don’t know your tools. They don’t know your CRM schema, your internal APIs, or the specific sequence of operations your business runs daily.
Fine-tuning changes that. A fine-tuned model learns your exact toolset — which tool to call, what arguments to pass, and when to chain multiple tools together. The result is an AI assistant that operates inside your actual workflow, not alongside it.
The techniques below are drawn from production fine-tuning work on Qwen 3.5, a current-generation open-source model, trained for multi-step tool calling across voice assistants and business applications.
Select the Right Base Model and Quantization
The model you start from determines everything that follows. Two decisions matter: which model family, and what precision to train at.
Model selection.Qwen 3.5 is currently the strongest open-source base for tool calling — it handles structured function calls more naturally than alternatives like Mistral 3. When evaluating a base model, prioritize native tool-calling ability over raw benchmark scores. A model that already understands function-call syntax requires less training data to become reliable.
Quantization.This is where most practitioners make their first costly mistake. The standard approach for reducing memory requirements — QLoRA 4-bit training — does not work reliably on Qwen 3.5. The quantization differences are too large. Full BF16 precision is required. Unsloth’s own documentation confirms this for both the mixture-of-experts and dense variants.
The practical impact: VRAM requirements jump from 12GB (where a consumer GPU suffices) to 75GB (where cloud hardware becomes necessary). Skipping this step produces a model that appears to train successfully but behaves unpredictably in production. There is no shortcut here.
Structure Training Data as Multi-Turn Conversations
Single-turn training examples — one user request, one tool call — produce models that work in demos and fail in production. The correct approach uses full multi-turn conversations that mirror real operating conditions.
Each training example should include:
- The complete system prompt with all tool definitions
- Previous conversation history — prior user messages and tool calls from the same session
- The current user request and the correct tool-calling sequence to fulfill it
A realistic training example looks like this: the user says “email Sarah the next F1 race.” The model must chain three tool calls — query the sports schedule API, look up Sarah’s email in contacts, then compose and send the email using Gmail. Each step feeds information to the next.
This approach inflates each training example from a few hundred tokens to roughly 7,000 tokens. That is the cost of production-realistic training. Accept it.
Store training data in JSON.One example per line in a JSONL file. The training script converts to whatever format the target model requires — Qwen, Mistral, or otherwise. This separates your data investment from any single model architecture.
Split Context from Training Signal
This is the single most important technique in the entire pipeline.Without it, the model learns to call tools on every interaction — even simple conversational exchanges or factual questions it could answer from its own knowledge.
The problem: when you include full conversation history in training examples and train on all of it, the model treats every message-response pair as a tool-calling pattern. Ask it the capital of Japan and it fires off a web search. Ask it to say good morning and it queries your calendar.
The correct approach: split each training example into context (not trained on) and signal (trained on).
- Context section: The full conversation history — previous user messages, previous tool calls, previous assistant responses. The model sees this but is not graded on it.
- Training signal: Only the most recent user request and the correct response. This is what the model learns from.
The model learns: “Given all this prior context, here is when and how to use a tool for this specific request.” It stops treating every interaction as a tool-calling opportunity. This distinction — context versus signal — is what separates a 60% accuracy model from a 99% accuracy model.
Use Cloud GPUs Instead of Consumer Hardware
A single training run costs $2–$12 on cloud hardware. A consumer GPU capable of the same work costs $3,000–$5,000 — if you can find one — and still cannot handle the 9-billion parameter model.
Cloud GPU marketplaces allow you to rent exactly the hardware you need for exactly the duration you need it:
- 2B parameter model: RTX 5090 (32GB VRAM), $0.43/hour, ~5 hours training, ~$2.14 total
- 9B parameter model: A100 (80GB VRAM), $1.13/hour, ~11 hours training, ~$12.43 total
When selecting an instance, prioritize network speed over marginal GPU cost savings.The difference between 500 Mbps and 6,000 Mbps upload speed is often less than $0.10/hour — but it determines whether downloading your trained model takes minutes or an hour. The few cents saved on a slower connection are lost many times over in rental time while waiting for file transfers.
Setup should be scripted, not manual. A setup script that accepts an instance ID and handles SSH connection, file transfer (system prompt, training data, training scripts), dependency installation, and environment validation eliminates the error-prone manual process and makes every training run reproducible.
Manage VRAM Through Batch Configuration
When training on limited VRAM, batch size is your primary control lever. Split the effective batch size across gradient accumulation steps — for example, a batch of 4 with 2 accumulation steps produces an effective batch of 8 while using roughly half the peak memory of a batch of 8 processed at once.
Set your maximum sequence length to match your actual training data.If your longest examples are 10,000 tokens, a sequence length cap of 8,000 silently truncates them — the model never sees the final tool calls in your longest examples. Measure your data, then set the cap.
Automate Completion and Cost Control
Training runs take hours. Watching a terminal for 5–7 hours is not a productive use of time, and forgetting about a completed run means paying for idle GPU rental.
The correct approach uses a polling script:
- Accept the instance ID, output directory, and estimated training duration
- Sleep for 80% of the estimated time (no point checking before then)
- Poll every 5 minutes for a completion marker file on the server
- On completion: download the LoRA adapter and exported model file
- Shut down the GPU instance automatically
This pattern means you can start a training run at 10 PM, go to sleep, and wake up to a downloaded model and a terminated instance. The polling script can run on any always-on machine — a local server, a small board computer, anything with SSH access.
The training script itself should write a simple completion marker (a text file) as its final action. The polling script checks for that file’s existence. Simple, reliable, no monitoring infrastructure required.
Export Both Formats
For LoRA fine-tuning on a model of this class, roughly 1% of total parameters become your fine-tuned adapter. Two epochs over the training data is a reasonable starting point — enough for the model to learn the patterns without overfitting to specific examples.
After training completes, export both the LoRA adapter (for flexible deployment) and a quantized GGUF file (for edge deployment on constrained hardware). The GGUF export can run on the same cloud instance before shutdown, adding minutes to the training run but saving a separate processing step later.
The Bottom Line
Fine-tuning a model for tool calling is a solved problem — but the details of execution determine whether the result is production-grade or a demo that falls apart under real conditions. The five techniques that matter most:
- Train at full BF16 precision on Qwen 3.5. No 4-bit shortcuts.
- Use multi-turn training examples with complete system prompts and tool definitions.
- Split context from training signal so the model learns when to call tools, not just how.
- Use cloud GPUs and script the setup for reproducibility.
- Automate polling and shutdown so training runs don’t waste money on idle hardware.
Total investment to reach a 99% tool-calling accuracy model: approximately $100 in cloud GPU credits across multiple iteration rounds. The pipeline, once built, makes each subsequent training run a $2–$12 operation with no manual intervention required.
In our One Weekend AI Masterclass workshops, we cover the practical side of AI model customization — including when fine-tuning makes sense for your business and when off-the-shelf tools are the better path. The goal is always the same: make AI work inside your workflow, not around it.