RecursionAI
Back to R&D
Past Project

Insurance AI Automation

A specialized 14B model that outperformed GPT-4o on owned hardware

Flagship case for specialized models. The client was paying ~$2,500/mo for a closed-API generalist and wanted to 4× usage. We LoRA fine-tuned Qwen2.5 14B for the specific task — it outperformed GPT-4o on their workload and gave them headroom to scale.

Specialized > Generalist

LoRA fine-tuned Qwen2.5 14B for the client's specific task. On their workload, the fine-tuned 14B model outperformed GPT-4o.

Self-hosted AI

Self-hosted Qwen2.5 14B for document parsing, analyses, and data extraction on ~48GB of hardware.

Scalability

Batching requirements sized for current workload and 4× scale — same hardware absorbs all of it.

Results

Eliminated a $2,500/mo recurring API cost, improved performance on a core feature, and 4×'d usage.

Tech Stack

TensorRT LLMOpenAI-CompatibleQwen2.5 14BPythonLoRA