Summary

Backbase published a study showing its 12-billion-parameter banking AI model outperformed GPT-4.1 in a live deployment at a large US financial institution. Over seven months, the model resolved 7.1 percentage points more customer queries across 3,297 interactions. The model was trained to refuse answers when evidence did not support a response (12% refusal rate vs 4.3% untuned vs 20.2% GPT-4.1). It scored 6.21 vs 5.72 on answer quality (10-point scale), with citation grounding improved by 2.3 points. Operating cost: ~$0.001/query (20-50x cheaper than GPT-4.1), training cost: ~$1,800. The research was led by Denys Katerenchuk (ex-Google, ex-IBM).

Key Facts

Why It Matters

Backbase's study challenges the assumption that bigger general-purpose models are always better for banking. The key insight: a smaller, domain-specific model trained to know its limits (higher refusal rate) can outperform a much larger general model on actual customer query resolution. The cost advantage (20-50x cheaper) makes this economically significant for banks processing millions of queries. The finding that data order matters more than data volume is a practical lesson for any financial institution building custom AI. This could accelerate the trend toward specialized banking AI models rather than relying on general-purpose LLMs.

Sources

Powered by Forestry.md