Insights/Technical Guide

How to Fine-Tune an AI Model for Tool Calling — The Right Way

Training a custom AI model to use your business tools costs $2-$12 per run. But the technique determines whether you get 99% accuracy or an assistant that fires off random tool calls. Here is what works.

KC

Kyle Cunningham

Founder & Lead Instructor, The Northline Institute

·10 min read

Why This Matters for Business

Off-the-shelf AI models are generalists. They can draft emails and summarize documents, but they don’t know your tools. They don’t know your CRM schema, your internal APIs, or the specific sequence of operations your business runs daily.

Fine-tuning changes that. A fine-tuned model learns your exact toolset — which tool to call, what arguments to pass, and when to chain multiple tools together. The result is an AI assistant that operates inside your actual workflow, not alongside it.

The techniques below are drawn from production fine-tuning work on Qwen 3.5, a current-generation open-source model, trained for multi-step tool calling across voice assistants and business applications.

Select the Right Base Model and Quantization

The model you start from determines everything that follows. Two decisions matter: which model family, and what precision to train at.

Model selection.Qwen 3.5 is currently the strongest open-source base for tool calling — it handles structured function calls more naturally than alternatives like Mistral 3. When evaluating a base model, prioritize native tool-calling ability over raw benchmark scores. A model that already understands function-call syntax requires less training data to become reliable.

Quantization.This is where most practitioners make their first costly mistake. The standard approach for reducing memory requirements — QLoRA 4-bit training — does not work reliably on Qwen 3.5. The quantization differences are too large. Full BF16 precision is required. Unsloth’s own documentation confirms this for both the mixture-of-experts and dense variants.

The practical impact: VRAM requirements jump from 12GB (where a consumer GPU suffices) to 75GB (where cloud hardware becomes necessary). Skipping this step produces a model that appears to train successfully but behaves unpredictably in production. There is no shortcut here.

Structure Training Data as Multi-Turn Conversations

Single-turn training examples — one user request, one tool call — produce models that work in demos and fail in production. The correct approach uses full multi-turn conversations that mirror real operating conditions.

Each training example should include:

  1. The complete system prompt with all tool definitions
  2. Previous conversation history — prior user messages and tool calls from the same session
  3. The current user request and the correct tool-calling sequence to fulfill it

A realistic training example looks like this: the user says “email Sarah the next F1 race.” The model must chain three tool calls — query the sports schedule API, look up Sarah’s email in contacts, then compose and send the email using Gmail. Each step feeds information to the next.

This approach inflates each training example from a few hundred tokens to roughly 7,000 tokens. That is the cost of production-realistic training. Accept it.

Store training data in JSON.One example per line in a JSONL file. The training script converts to whatever format the target model requires — Qwen, Mistral, or otherwise. This separates your data investment from any single model architecture.

Split Context from Training Signal

This is the single most important technique in the entire pipeline.Without it, the model learns to call tools on every interaction — even simple conversational exchanges or factual questions it could answer from its own knowledge.

The problem: when you include full conversation history in training examples and train on all of it, the model treats every message-response pair as a tool-calling pattern. Ask it the capital of Japan and it fires off a web search. Ask it to say good morning and it queries your calendar.

The correct approach: split each training example into context (not trained on) and signal (trained on).

  • Context section: The full conversation history — previous user messages, previous tool calls, previous assistant responses. The model sees this but is not graded on it.
  • Training signal: Only the most recent user request and the correct response. This is what the model learns from.

The model learns: “Given all this prior context, here is when and how to use a tool for this specific request.” It stops treating every interaction as a tool-calling opportunity. This distinction — context versus signal — is what separates a 60% accuracy model from a 99% accuracy model.

Use Cloud GPUs Instead of Consumer Hardware

A single training run costs $2–$12 on cloud hardware. A consumer GPU capable of the same work costs $3,000–$5,000 — if you can find one — and still cannot handle the 9-billion parameter model.

Cloud GPU marketplaces allow you to rent exactly the hardware you need for exactly the duration you need it:

  • 2B parameter model: RTX 5090 (32GB VRAM), $0.43/hour, ~5 hours training, ~$2.14 total
  • 9B parameter model: A100 (80GB VRAM), $1.13/hour, ~11 hours training, ~$12.43 total

When selecting an instance, prioritize network speed over marginal GPU cost savings.The difference between 500 Mbps and 6,000 Mbps upload speed is often less than $0.10/hour — but it determines whether downloading your trained model takes minutes or an hour. The few cents saved on a slower connection are lost many times over in rental time while waiting for file transfers.

Setup should be scripted, not manual. A setup script that accepts an instance ID and handles SSH connection, file transfer (system prompt, training data, training scripts), dependency installation, and environment validation eliminates the error-prone manual process and makes every training run reproducible.

Manage VRAM Through Batch Configuration

When training on limited VRAM, batch size is your primary control lever. Split the effective batch size across gradient accumulation steps — for example, a batch of 4 with 2 accumulation steps produces an effective batch of 8 while using roughly half the peak memory of a batch of 8 processed at once.

Set your maximum sequence length to match your actual training data.If your longest examples are 10,000 tokens, a sequence length cap of 8,000 silently truncates them — the model never sees the final tool calls in your longest examples. Measure your data, then set the cap.

Automate Completion and Cost Control

Training runs take hours. Watching a terminal for 5–7 hours is not a productive use of time, and forgetting about a completed run means paying for idle GPU rental.

The correct approach uses a polling script:

  1. Accept the instance ID, output directory, and estimated training duration
  2. Sleep for 80% of the estimated time (no point checking before then)
  3. Poll every 5 minutes for a completion marker file on the server
  4. On completion: download the LoRA adapter and exported model file
  5. Shut down the GPU instance automatically

This pattern means you can start a training run at 10 PM, go to sleep, and wake up to a downloaded model and a terminated instance. The polling script can run on any always-on machine — a local server, a small board computer, anything with SSH access.

The training script itself should write a simple completion marker (a text file) as its final action. The polling script checks for that file’s existence. Simple, reliable, no monitoring infrastructure required.

Export Both Formats

For LoRA fine-tuning on a model of this class, roughly 1% of total parameters become your fine-tuned adapter. Two epochs over the training data is a reasonable starting point — enough for the model to learn the patterns without overfitting to specific examples.

After training completes, export both the LoRA adapter (for flexible deployment) and a quantized GGUF file (for edge deployment on constrained hardware). The GGUF export can run on the same cloud instance before shutdown, adding minutes to the training run but saving a separate processing step later.

The Bottom Line

Fine-tuning a model for tool calling is a solved problem — but the details of execution determine whether the result is production-grade or a demo that falls apart under real conditions. The five techniques that matter most:

  1. Train at full BF16 precision on Qwen 3.5. No 4-bit shortcuts.
  2. Use multi-turn training examples with complete system prompts and tool definitions.
  3. Split context from training signal so the model learns when to call tools, not just how.
  4. Use cloud GPUs and script the setup for reproducibility.
  5. Automate polling and shutdown so training runs don’t waste money on idle hardware.

Total investment to reach a 99% tool-calling accuracy model: approximately $100 in cloud GPU credits across multiple iteration rounds. The pipeline, once built, makes each subsequent training run a $2–$12 operation with no manual intervention required.

In our One Weekend AI Masterclass workshops, we cover the practical side of AI model customization — including when fine-tuning makes sense for your business and when off-the-shelf tools are the better path. The goal is always the same: make AI work inside your workflow, not around it.

Frequently Asked Questions

What is AI fine-tuning for tool calling?

Fine-tuning for tool calling is the process of training an AI model to use your specific business tools — your CRM, email system, scheduling platform, or any API — by showing it examples of when and how to call each tool. The result is an AI assistant that operates inside your actual workflow rather than alongside it. A fine-tuned model learns which tool to use, what arguments to pass, and when to chain multiple tools together to complete a request.

How much does it cost to fine-tune an AI model?

A single training run costs $2-$12 on cloud GPU hardware depending on model size. A 2-billion parameter model trains in about 5 hours on a cloud GPU at $0.43/hour (roughly $2.14 total). A 9-billion parameter model takes about 11 hours on a larger GPU at $1.13/hour (roughly $12.43 total). Including trial-and-error iterations, expect to spend approximately $100 to develop a reliable pipeline — far less than the $3,000-$5,000 cost of a consumer GPU that still cannot handle larger models.

What is the most important technique for training AI to call tools?

The single most important technique is splitting context from training signal. Each training example should include full conversation history (which the model sees but is not graded on) and the most recent user request with the correct tool-calling response (which the model is trained on). Without this split, models learn to call tools on every interaction — even simple questions they could answer from their own knowledge. With it, models achieve 99% accuracy on knowing when to use tools versus when to respond conversationally.

Do I need expensive hardware to fine-tune AI models?

No. Cloud GPU marketplaces allow you to rent exactly the hardware you need for exactly the duration you need it. A 2-billion parameter model requires 32GB of VRAM (available for $0.43/hour), while a 9-billion parameter model requires 80GB (available for $1.13/hour). Consumer GPUs max out at 32GB and cost $3,000-$5,000 to purchase. Cloud rental is more cost-effective for all but the highest-volume training operations.

Ready to Get Started?

Stop reading about AI. Start using it.

One Weekend AI Masterclass is a rigorous 2-day in-person workshop that takes business leaders from AI-curious to AI-competent. 25 seats per city.