MsgHub Blog

Notes from the engine room.

Engineering deep-dives, India CPaaS industry updates, AI cost optimisation, and the occasional unflattering postmortem.

engineering 14 May 2026 Surapura

How we built RAG on 397B params for $0.005 per conversation

The stack and math behind running a production WhatsApp AI on Qwen3.5 397B, Llama 4 Maverick and OpenAI embeddings — at one-twentieth the cost of a GPT-4 baseline.

#rag#llm-cost#qwen#whatsapp-ai#production