Summary

Backbase's study demonstrates that a smaller, domain-specific AI model (12B parameters) can outperform a much larger general-purpose model (GPT-4.1) in live banking deployment. The key insight: training the model to refuse answers when evidence is insufficient (12% refusal rate) actually improved customer query resolution by 7.1 percentage points. The model was 20-50x cheaper to run (~$0.001/query) and cost only ~$1,800 to train. The research also found that data order matters more than data volume — training sequentially (general finance → calibrated refusal) produced 40%+ better results than combining all data at once. The work was led by Denys Katerenchuk (ex-Google, ex-IBM) through Backbase AI Research (acquired via Kasisto).

Key Points

Sources

Powered by Forestry.md