Local AI Platform
Fine-tuning and retrieval-augmented generation, engineered in-house and run entirely on our own hardware — no cloud model, no per-token bill, no document ever leaving the building.
Two Ways We Ship AI
We run two AI systems entirely on our own infrastructure. One fine-tunes a small open model on a client's own product data until it answers support questions correctly. The other grounds a model in a specific document — a financial report, a filing — so it answers only from what's actually on the page.
Both run on a single consumer GPU, with nothing sent to a third-party model provider.
Wire up a general-purpose cloud model, prompt it with a system message, and hope it doesn't invent a feature, a page or a policy that doesn't exist — while every question leaves the building and shows up on a bill.
Fine-tune a 3-billion-parameter open model on the product's real data, or ground it in the real document, so an answer is either learned from the truth or read directly off it. Inference runs on hardware we own.
XponentShift Limited — the same in-house engineering team that builds our client products builds the AI layer that runs inside them: dataset generation, fine-tuning, retrieval engineering and evaluation, end to end, with no outsourced model vendor in the loop.
How This Differs From
a Bolted-On Chatbot
What We
Actually Built
Six pieces of engineering, each solving a specific, measured failure — not generic AI-integration boilerplate.
Dataset Engineering
2,312 synthetic training examples generated programmatically from the live product, its database schema, and its real navigation, pricing and status values — never from generic knowledge of how the product works. An independent validator rejects the build on any contradiction of a pinned product fact, any cross-split data leakage, or any invented navigation path.
Fine-Tuning Pipeline
Supervised fine-tuning on Qwen2.5-3B-Instruct, trained on a free Google Colab T4 GPU end to end, merged and exported to GGUF via llama.cpp, and imported into Ollama behind an explicit chat template — small enough to run on a laptop GPU, tuned enough to know the product cold.
Document Ingestion
A custom pdfplumber pipeline extracts prose and tables page by page, repairing rendering artifacts most pipelines miss entirely: double-struck bold text, letter-spaced section headings with no ruling line, and decorative panels that parse as "tables" but hold no data.
Retrieval Engineering
bge-m3 embeddings (1024-dimension, 8,192-token context) stored in Postgres via pgvector, so a dense financial table is one vector instead of being split mid-grid. Retrieval blends cosine similarity with lexical term overlap and reserves dedicated slots for tables.
Conversational Memory
"And is that a monthly charge?" is rewritten into a self-contained question before retrieval ever runs, with deterministic safety checks that reject any rewrite inventing a period, company or figure nobody in the conversation actually said.
Evaluation & Infrastructure
A 260-question held-out benchmark plus a separate 52-question "real feature" set, both rule-graded and exported to a reviewable Excel workbook. Deployed on Docker with NVIDIA GPU passthrough and exposed for stakeholder review over a Cloudflare Tunnel — no cloud hosting bill required.
Four Problems That Don't
Show Up in a Demo
Reading the Wrong Column
A financial table spanning five quarters put the right line item in front of the model — with the wrong period's figure sitting right next to it. Asked for Q2-2026 Adjusted EBITDA, it returned the year-ago number because both columns shared the label "Q2," a pattern traced to 15 of 44 evaluation errors. The fix rewrites every table so each figure states its own line item and period inline — "Adjusted EBITDA: Q2-2026 = 3,273" — before the model ever reads it, removing the column arithmetic entirely.
Follow-Up Questions With No Noun in Them
"And net income?" carries no company, no period and no subject of its own — embedded and searched literally, it retrieves nothing useful and can route an entire conversation to the wrong document. The fix rewrites every dependent-sounding message into a standalone question using the conversation's own established subject and period, then checks it against two hard rules: nothing in the rewrite may be a number nobody said, and nothing the user asked about may be dropped.
A Benchmark That Rewarded Saying No
A third of the held-out evaluation set scored the model on correctly withholding information, so a fine-tune that simply got more cautious scored higher on paper — while a real support conversation, full of ordinary product questions, is exactly the slice that benchmark barely tested. The fix built a second evaluation set of real, answerable product questions specifically to catch a model that scores well by refusing more; a release decision reads both sets together, never either alone.
Four Gigabytes of GPU Memory
A side-by-side base-vs-fine-tuned comparison tool needs two 3-billion-parameter models loaded at once — which don't both fit in VRAM on a consumer GPU. The fix pins Ollama to one loaded model at a time; the second request queues and waits for the swap rather than crashing the container, so the whole platform, including a live model comparison, runs on a single laptop GPU.
From Raw Data to a
Deployed, Evaluated Model
Baseline
Measure how the untuned base model performs on the actual target task before writing a single training example — so an improvement claim is measured against a real number, not assumed.
Dataset
Generate a synthetic supervised fine-tuning dataset straight from the product's own code, schema and copy. Validate it for contradictions, cross-split leakage and any real personal data before it touches a training run.
Train
QLoRA fine-tuning on a free-tier GPU, checkpointed regularly and synced off the training VM as it runs, so an interrupted session resumes instead of losing hours of compute.
Evaluate
Score the fine-tune against the base model on a held-out set using machine-checkable rules, plus a second set built to catch over-caution. Export every answer to a workbook a human can actually read.
Deploy
Import the trained model into Ollama behind an explicit chat template, serve it through our Next.js stack on Docker with GPU passthrough, and expose it for review over a Cloudflare Tunnel — no cloud GPU bill.
Where This Goes Next
- Second training round covering prompt-injection handling and stale-state awareness
- Onboarding-flow coverage, the largest remaining gap in real product questions
- Regrade tooling so a grading-rule fix never requires re-running the model
- Multi-document, multi-company retrieval at scale, beyond a single ingested report
- Automated regression evaluation on every dataset change, not just before a release
- A larger local hardware tier for concurrent model serving
- Fine-tuning-as-a-service: the same pipeline applied to any client's own product data
- Packaged on-prem deployments for clients with strict data-residency or GDPR requirements
- Routing between a small tuned model and a larger reasoning model by query difficulty
// Where fine-tuning meets retrieval — and neither one leaves the building
Want AI like this
built for your business?
Designed and built entirely by XponentShift Limited — the same in-house team, the same standard of measurement, applied to your product's own data.