The short version
Anthropic's Claude is the primary AI provider across the Mirembe Muse platform. It runs Nova, the companion inside VarsityOS. It runs the agents inside AdminOS. It runs StokvelOS, the Sanyu assistant, our internal JarvisOS, and the Mirembe AI assistant you can talk to on this very site.
One product does not use Claude. K53 Drill Master, our learner's-licence prep app, runs its AI Tutor on OpenAI's gpt-4o-mini. That is a decision, not an accident. We will get to why.
Why Claude is the default
Three reasons, in order of how much they matter to us.
First, the quality of reasoning. Most of what our products do is not trivia lookup — it is following instructions carefully, holding context, and not going off the rails when a user asks something unexpected. Claude is consistently strong here.
Second, safety. A lot of our users are students, first-time founders, and people managing other people's money in a stokvel. The default behaviour of the model matters when the stakes are real.
Third — and this one is unglamorous but decisive — prompt caching. It is what makes running AI affordable at African scale, which we will come back to, because it shapes almost everything.
Matching the model to the task
Having a default provider does not mean using one model for everything. Within Claude, we match the model to the job.
For high-volume, low-latency work — chat, streaming replies, the assistant that has to feel instant — we use a fast, cheap model like Claude Haiku. For heavier reasoning, where a wrong answer costs more than a few extra cents, we step up to Sonnet or Opus.
The rule we hold ourselves to is simple: never pay Opus prices for a Haiku job. A studio that runs its most expensive model on every request is either not watching its costs or passing them to you.
The cost constraint is a design constraint
Building AI products in South Africa means the maths has to work in rands, for users who are price-sensitive by necessity. A model that is brilliant but too expensive per request is, for our market, the wrong model.
This is where prompt caching earns its place. A large part of most prompts — the instructions, the product context, the rules — is the same every time. Caching that shared portion means we are not paying to re-process it on every message. At the volumes VarsityOS or AdminOS run, that is the difference between a feature that ships and one that stays a demo.
So when we say we chose a model, we mean we chose the whole equation: reasoning quality, latency, safety, and cost per request at the scale the product actually runs.
The honest exception: K53 Drill Master
K53 Drill Master uses OpenAI's gpt-4o-mini for its AI Tutor. We could have forced Claude in to keep a tidy story. We did not.
Part of it is range — we want our team fluent across providers, not married to one vendor's way of doing things. Part of it is that gpt-4o-mini was simply the right fit for that product, at that price, doing that job.
We will never tell you K53 runs on Claude. It does not. Stating that plainly costs us nothing and tells you something more useful than any marketing line: we know exactly what is under the bonnet of each thing we build, and we will tell you the truth even when the truth is slightly off-brand.
The real point
Here is the thing most AI conversations get backwards. The model is rarely the hard part. The hard part is the system around it — the prompts, the caching, the guardrails, the retries, the way it fails gracefully when the network drops in an area with bad signal, the way it hands off to a human when it should.
Swapping one strong model for another rarely changes a product's fate. Good system design almost always does.
So we treat honesty about our tools as a feature, not a confession. A studio that will tell you exactly which model runs which product — including the one that isn't their default — is a studio you can trust with the parts you cannot see. That is the whole point of telling you.
.jpg&w=3840&q=75)
.jpg&w=3840&q=75)
.jpg&w=3840&q=75)