AssociationAI / AI Literacy
Trihelix AI team Published

Tutorial

Advanced

Fine-Tune Your Own Model with Unsloth

Train an open model on your association's own Q&As with Unsloth's QLoRA workflow: your voice, your data rules, member-facing or internal.

Time needed: 4 to 6 hours of staff time to curate the dataset (the long pole), plus a 1 to 3 hour training run

Before you start:

  • Fifty or more published Q&As, FAQs, or helpdesk answers your association could publish as-is
  • A Google account for the free Colab GPU, or a PC with an NVIDIA GPU and IT help
  • One reviewer who reads every training pair before it trains

Non Dues Revenue Fine Tuning Member Service

By the end of this tutorial you will have a language model that answers in your association’s voice, fine-tuned on your own Q&A pairs with Unsloth’s free toolkit on a free cloud GPU or your own machine.

Try the finished asset first: our training-data studio builds the exact file a fine-tune reads. Write Q&A pairs, watch them format into JSONL, and run the PII scan. The starter kit (notebook, converter, template, checklists) is free with your email.

The first decision is the only one with real risk: what data may train. Colab’s free GPU is Google’s computer, and anything you upload there leaves your building.

The eight steps

  1. Decide what may train. Apply the publish test: unpublished text does not go to Colab. Past FAQs, event copy, and member questions with names stripped out pass. Raw helpdesk exports fail. Models memorize training text: Carlini and colleagues extracted hundreds of verbatim text sequences, including names and phone numbers, from a model’s training data, some appearing in just one document. Pick your route: Colab’s free T4 for material that passes the publish test, a local NVIDIA GPU for anything sensitive, or dataset now and hardware later.

A decision diagram for what training data may go to Colab versus staying on a local GPU.

  1. Gather your raw material. Pull from what you already publish: the FAQ page, the certification handbook, and event instructions. Add past member questions only after stripping names, companies, and contact details. The next step turns them into pairs.

  2. Write the Q&A pairs, and review every one. Turn each item into a real member question and your best staffer’s answer. An AI tool may draft the pairs, but a human approves every one, because the model repeats your phrasing back, mistakes included. Zhou and colleagues found that only limited instruction tuning data is necessary for high quality output, with their LIMA model’s responses equivalent or preferred to GPT-4 in 43% of cases. Fifty careful pairs beat five hundred sloppy ones. Fine-tuning teaches voice, not facts: Gudibande and colleagues found that imitation models are adept at mimicking ChatGPT’s style but not its factuality.

  3. Format the pairs and scan them in the training-data studio. Open the training-data studio, add your pairs: it formats them into JSONL, one object per line with instruction, input, and output fields. Run the built-in PII scan for email addresses, phone numbers, and ID patterns. The scan catches patterns, not names; the human review from step 3 still stands. For the general redaction habit, see our tutorial on what member data is safe to paste into AI. Download the resulting training.jsonl.

A labeled example of one JSONL training object with instruction, input, and output fields.

  1. Open the starter notebook and load the base model. Upload the kit’s notebook to Colab and your training.jsonl when it asks, then run the cells. The notebook installs Unsloth, prints your VRAM, and loads an 8B-class open model in 4-bit. Dettmers and colleagues fine-tuned a 65B parameter model on a single 48GB GPU with QLoRA. At our scale, Unsloth’s requirements table puts an 8B QLoRA fine-tune at 6 GB VRAM minimum, and Unsloth’s Colab guide confirms the free tier provides a T4. Video memory is the number that matters, not system RAM.

  2. Attach the adapters and start the training run. The notebook freezes the base model and trains only the small QLoRA adapters, which is why a free GPU suffices. Smoke-test with sixty steps, then train two epochs and watch the loss fall and flatten. If the loss climbs, stop: the usual cause is a malformed JSONL line.

A three-box flow from frozen base model through training adapters to a merged model.

  1. Run the ten-question scorecard. Hold out ten member questions. Ask each to the base model and your fine-tuned model, then score both with the kit’s rubric. The full protocol sits in the check section below.

  2. Export the weights and pick the deploy route. Save the small adapter files from the notebook; they are the fine-tune. For a member-facing helper, merge the adapters and serve the model behind your member login as a new reason to belong. For internal work, point the model at the repetitive queue: certification questions, billing questions, event logistics. Fine-tuning teaches the model how your association sounds; if you need it to answer from specific documents instead, that is a retrieval job, and our tutorial on running Laya triage on your own machine shows the local-inference route.

Three deployment routes: Colab, local GPU, or dataset now with training later.

One association’s training set

This is a teaching example, not a case study. The fictional Harborlight Marina Association fields the same dockmaster certification questions every spring. Its education manager de-identifies past member questions into one hundred twenty pairs; every pair passes the publish test, so she trains on Colab. Example pair: “When does my dockmaster certification expire?” paired with “March 1 each year; renew in the member portal and your new card arrives by mail.” The member-facing result answers new certification questions on the member portal in the same plain Harborlight phrasing. Internally, the same model drafts replies to certification emails for staff review. Example pairs and outputs, not real member text.

Check it: the ten-question scorecard

  1. Write ten questions members actually ask, none appearing in training.jsonl.
  2. Ask each to the base model and to the fine-tuned model, and record both answers.
  3. Score each answer 0 to 2 on correctness, voice, and safety: no invented policy, no PII, no confident guesses.
  4. Pass bar: the fine-tuned model outscores the base on at least seven of ten, and never scores 0 on safety.
  5. A safety 0 anywhere stops the ship. Fix the training pairs, retrain, re-score.

Five ways to waste a training run

Training raw PII on Colab. See the memorization finding in step 1. De-identify or go local.

Ten pairs and hope. A tiny dataset teaches recitation, not voice. Fifty reviewed pairs is the floor.

No held-out test. Training questions always look good. The ten unseen questions are the honest grade.

Expecting new facts. Fine-tuning teaches phrasing. Fine-tuning will not teach it your new dues schedule from three examples; put new facts in the pairs verbatim.

Skipping the base comparison. Without the base model’s answers beside yours, every improvement is a feeling. Score it.

Sources (6)