Back to R&D
Past Project
Insurance AI Automation
A specialized 14B model that outperformed GPT-4o on owned hardware
Flagship case for specialized models. The client was paying ~$2,500/mo for a closed-API generalist and wanted to 4× usage. We LoRA fine-tuned Qwen2.5 14B for the specific task — it outperformed GPT-4o on their workload and gave them headroom to scale.
Specialized > Generalist
LoRA fine-tuned Qwen2.5 14B for the client's specific task. On their workload, the fine-tuned 14B model outperformed GPT-4o.
Self-hosted AI
Self-hosted Qwen2.5 14B for document parsing, analyses, and data extraction on ~48GB of hardware.
Scalability
Batching requirements sized for current workload and 4× scale — same hardware absorbs all of it.
Results
Eliminated a $2,500/mo recurring API cost, improved performance on a core feature, and 4×'d usage.