
maainstream.com — the offer, stated in the client's own terms.
Online programs are sold on a promise and settled on a completion rate. Someone pays for a challenge, a course, or a mentorship, opens it once, hits the first thing they do not understand at 11pm, and never comes back. Nobody is there. The refund arrives three weeks later.
The usual fix is headcount: hire support staff, hire community managers, hire coaches. It works, and it stops working the moment volume grows faster than the payroll. Maainstream is the other answer — an AI coach that lives inside the program, knows the material, knows where each customer actually is, and answers at 11pm.
I built it end to end: the product, the AI system underneath it, and the infrastructure it runs on. Solo, from the first commit in December 2025 to a platform carrying six figures of production conversations.
What the thing actually is#
Customers do not talk to a chatbot bolted onto a marketing site. They log into the platform their program lives on and talk to a coach that has the program in front of it and their own history behind it.

One client's coach. Every instance is the same engine under different branding, program content, starter prompts, and calls to action.
That framing did most of the architectural work. A support widget is stateless and generic — it answers the question in front of it. A coach is neither. It has to know that this person is on day 4 of a 21-day challenge, that they skipped the module the current question depends on, and that they said last week they were about to give up. The conversation is a side effect of the state, not the other way around.
It is also multi-tenant from the first migration. An organisation owns agents; an agent owns its persona, knowledge base, learning stages, suggested prompts, and sidebar banner. Every client configures their own coach without any of it being a fork. One engine, many coaches.
The backend is Express and TypeScript over PostgreSQL, with Kysely rather than an ORM — the queries that matter here are vector searches with relational filters bolted to them, and hand-written SQL with generated types beats fighting a query builder that wants to hide the SQL. The frontend is React 19 on TanStack Router with SSR, streaming chat over SSE.
The AI system#
One agent, a small set of tools, a hard step limit#
The conversational layer is deliberately not a swarm. It is a single agent, built on the Vercel AI SDK, with two tools — search_knowledge and save_lead_info — and stopWhen: stepCountIs(3).
That step cap is the important part. An uncapped agent loop against a paying customer's chat window is an unbounded bill and an unbounded latency budget, and in practice the third step is where useful work stops and pacing back and forth begins. Three steps is enough to search, read, and answer.
The specialisation lives elsewhere. Rather than routing a live conversation between sub-agents, the system runs a set of narrow, single-purpose LLM jobs around the conversation — memory extraction, conversation summarisation, mastery assessment, lead-field extraction. Each has one schema, one prompt, and one definition of correct. They run on a queue, off the response path, where a slow or failed call costs nobody a wait.
This is the trade I would make again. In-conversation routing adds a hop and a failure mode to the one interaction the customer actually experiences; moving the specialists to the background keeps the live path short and makes each specialist independently testable.
Retrieval, scoped to where the learner actually is#
The coach is only credible if it answers from the material the customer paid for. Each client's program is ingested, chunked, embedded, and retrieved at answer time.
Chunking splits markdown on its headings and keeps the heading attached to the chunk as context, so a fragment retrieved out of a long module still carries what section it came from.
Retrieval is a tool the agent chooses to call, not a blind pre-fetch stapled to every turn. Most messages in a coaching conversation — "ok thanks", "I'll try tonight" — need no retrieval at all, and paying an embedding call plus a vector search on each of them is pure waste at six figures of conversations.
More importantly, retrieval is scoped by stage. The search is filtered to the part of the program the learner has actually reached, so the coach cannot answer a day-3 question with day-19 material and spoil the sequence the client designed.
That filter is why the vectors live in PostgreSQL through pgvector rather than in a dedicated vector database. The predicates that matter — this org, this agent, this stage, this session — are relational. Keeping embeddings in the same database as that state makes them a WHERE clause instead of a distributed join across two systems held together by hope.
Memory: hybrid search, with decay#
Every ten new messages, one background job both summarises the conversation and extracts what the user has actually stated about themselves into five categories — preference, fact, goal, context, behaviour — each with an importance score, under a prompt whose first instruction is to never speculate. Inference is how a memory store fills up with confident nonsense. A separate sweep catches conversations that go idle before the next interval and gives them a final pass, so someone who stops mid-thread does not lose what they said.
Retrieval of those memories is hybrid and non-fatal: a pgvector cosine search and a keyword search run against the same store, their results merge, and scores decay with age so a preference stated three months ago does not outrank one from yesterday. At most five memories reach the system prompt. If the vector half fails, the keyword half still answers — an embedding provider having a bad afternoon degrades memory rather than breaking chat.
Mastery: the part that makes it a coach#
This is what separates the product from a support bot. Each stage of a program defines mastery topics, and after each exchange a background job assesses whether the learner demonstrated understanding of any of them.
Two design decisions carry it:
The assessor only sees topics that are plausibly in play. An embedding pass pre-filters the stage's topics by similarity to the recent exchange, keeping at most five above threshold. Sending an LLM forty topics and asking which ones were demonstrated produces expensive mush; sending it five produces a judgement.
Doing counts, not just explaining. The assessment prompt treats "I created my account" or "it didn't work, I need to redo it" as evidence of understanding, alongside correct explanations. In a challenge, the demonstration of mastery is usually an action reported in passing, not an essay. Mastery is granted on a confidence score weighted between history and recency, against a threshold the client sets per topic.
The whole path is cost-shaped: assessment is skipped entirely for messages under twenty characters, history is capped at three exchanges, and user messages are truncated. None of that is visible to the customer, and all of it is the difference between a feature and a line item.
One gateway, three capabilities, cross-vendor fallback#
Every call — chat, utilities, embeddings — goes through OpenRouter on a single key per organisation. One provider to configure, one bill to read, and the full live model catalogue available to swap from the admin UI.
Capabilities are separated rather than models being picked once for the whole system:
- chat — the turn the customer reads. Defaults to Claude Sonnet.
- utilities — summarisation, memory extraction, mastery assessment, lead extraction. Defaults to a Gemini Flash-class model.
- embedding — pinned to 1536 dimensions, because that is the width of the
vector(1536)columns; changing it means a migration that re-embeds everything.
Each capability carries a fallback chain, and the chains are deliberately spread across different upstream vendors, so one vendor's outage cannot take chat down. On a rate-limit or quota error the agent retries down the chain — but only if response headers have not been sent yet, because there is no honest way to restart a stream a customer is already reading. A failed provider also raises an alert on the organisation, so the client finds out from the product rather than from a customer.
Holding the line on prompt injection#
An AI coach that can be talked out of its instructions is a liability on someone else's brand. There is a dedicated security layer on both sides of the model:
- Input is scored against a catalogue of injection patterns — authority spoofing ("provided to me by Anthropic"), role hijacking, system-prompt extraction — with risk scores driving allow, flag, or block, and repeat offenders getting their session marked suspicious and eventually cut off.
- Retrieved content is escaped before it reaches the prompt. The knowledge base is client-uploaded, which makes it an injection surface like any other user input.
- Output is validated for system-prompt leakage, role confusion, and structural tag leaks before it reaches the customer, with a fallback response if it fails.
All of it is unit-tested, which is what you want from the layer whose failure mode is a screenshot on social media.
The unglamorous half#
Ingestion, embedding, memory processing, summarisation, and mastery assessment run as BullMQ jobs on Redis with retries and backoff, not inside request handlers. A customer's message should not wait on a client uploading their back catalogue. Mastery and stage advancement push back to the browser over SSE via Redis pub/sub, so a learner sees a topic tick over without polling.
The whole thing runs on Railway, with Pino to Loki and Prometheus to Grafana behind it. The deployment story is boring on purpose, and it should stay that way for a long time.
What the operator sees#
An AI product nobody can inspect is a liability. Token usage and the exact cost OpenRouter billed are recorded per call, tagged with agent, conversation, session, and user — which is what makes the dashboard real numbers instead of estimates.

Conversation volume, message distribution, token spend, and lead capture, per client and per period.
Two things earn their place there. Messages per conversation is the distribution that tells you whether the coach is working: a bar at 0-1 is a product nobody engages with, and the mass sitting in the 21-50 band is the shape of a customer who keeps coming back. And lead capture is a first-class object rather than a transcript to be mined later — the pre-sale half of the product exists to turn a finished challenge into a qualified conversation, so qualification has to be structured data the moment it happens.
The lead-capture flow is the one place the system deliberately writes past the model: once the agent's stream finishes, if required fields are still missing, a single utility call both extracts what the last message revealed and drafts the follow-up question, which is appended to the stream the customer is already reading. One extra call, no second round trip, and the question arrives attached to a real answer rather than as an interrogation.
Results#
The numbers below are what the platform has measured in production, and what the client publishes publicly:
- 100,000+ conversations handled across client programs.
- 45-60% program completion, against the 5-20% that this market treats as normal. More customers reaching the end of a challenge also means more of them reaching the offer at the end of it.
- Refund rate divided by three. People who finish do not ask for their money back.
- 2.4 full-time support roles' worth of work absorbed in 11 weeks on a single client, with a 269:1 ratio of positive to negative messages.
- On one €14M challenge, 1,000+ spontaneous applications to a mentorship starting at €3,000 — a downstream effect of more people finishing the free thing first.
The engineering point in all of that: none of these numbers came from a better model. They came from the coach knowing who it was talking to, and from the boring work of keeping that knowledge cheap, scoped, and current.
My goal: turn what you need into software that is reliable, maintainable, and genuinely usable in production.
Email me about your project