How we built RAG on 397B params for $0.005 per conversation
The stack and math behind running a production WhatsApp AI on Qwen3.5 397B, Llama 4 Maverick and OpenAI embeddings — at one-twentieth the cost of a GPT-4 baseline.
Engineering deep-dives, India CPaaS industry updates, AI cost optimisation, and the occasional unflattering postmortem.
The stack and math behind running a production WhatsApp AI on Qwen3.5 397B, Llama 4 Maverick and OpenAI embeddings — at one-twentieth the cost of a GPT-4 baseline.