Desktop app · Local first
Pick specialists, run each of them on whatever model you like — on your own hardware or hosted — and give them a shared memory of your work. Everything stays on your machine.
macOS · Windows · Linux — early access opening soon
A normal chat window gives you one voice, on one company's model, with no memory of your work.
Run one role on a model on your own hardware, another on a hosted frontier model, and compare them side by side in the same conversation.
A Security Expert, a QA Engineer, a Copywriter — each set up for its job, each on the model that suits it.
Send a question to three roles and let them build on each other's answers, each one addressing the part it's best at.
Your documents and past conversations stay searchable and get used automatically, instead of being re-explained every session.
Everything grouped by what you're working on, so three jobs at once never bleed into each other.
Conversations, memory and API keys live in a database on your computer. Nothing is uploaded to run the app.
Everywhere else you pick a model and that's the product. Here it's a choice you make per role, per task, per project.
Ollama and MLX, pointed at whatever you're running locally. No cost per request, nothing leaving the machine.
OpenAI, Anthropic, Google — the frontier models, for the work that genuinely needs them.
Azure, AWS Bedrock, vLLM, or any endpoint you can give a URL to, with custom auth headers where you need them.
Put the cheap fast model on drafting and summarising. Put the expensive one on the architecture review. Keep anything sensitive on a model that never leaves your machine, while the rest of the team runs hosted. One conversation, several models, no juggling of tools or tabs.
The Prompt Analyzer runs the same prompt across several models side by side, with what each result cost.
Most of us over-buy on model size out of caution. Run the comparison against your real work and a much cheaper model usually handles the bulk of it — and you learn exactly which tasks are worth the expensive one.
One searchable place where everything you've given the app lives.
Drop in specs, contracts, policies, notes, reports, source code — whatever the work involves. It gets indexed and becomes background knowledge the AI actually draws on, rather than something you paste in again every time.
Your conversations are part of it too. Every session is kept and searchable, labelled with which role said what, so a decision made three weeks ago in a team discussion is still findable.
An answer draws on everything the person is entitled to and nothing beyond it — a shared base for the group they belong to, the project they're in, and the session they're working in right now.
Standards, policies, reference material and house templates that everyone in the group works from.
The documents, specs and decisions belonging to one job — kept apart from every other job.
What you and the team have said so far, retained and searchable long after you close it.
Memory is bound to projects and groups, so one client's material never turns up in an answer about another client's work.
Every document added makes every role using it more useful — and in a shared base memory, it does that for everyone at once.
Four modes, each adding to the last. Start wherever you are and move up when the work asks for it — one person, a small team, or an organisation.
A chat box, a model picker, light or dark. Everything else stays out of the way until you want it.
The role library, pre-built teams, the template library, simple step-by-step workflows, and cost tracking so you can see what you are spending.
The visual workflow designer with branching, memory and document indexing, template authoring, MCP tooling, the Prompt Analyzer, and detailed logging when something misbehaves.
Audit logging, budget controls, a governance dashboard, compliance workflows, and executive roles — for work that has to be answerable to someone.
Not everyone should get the same setup — or see the same memory. Roles decide what a person can do; groups decide what they can draw on.
Create a group and give it a base memory: the standards, reference documents, policies and house templates everyone in that group works from. Anyone in the group draws on it automatically, from their first session. Update the base once and everybody's answers improve at the same time — nobody has to be told to go and re-read anything.
Groups can overlap. Someone on a spatial project can sit in both the GIS group and the delivery group, and get both bases without either one leaking to people outside it.
What a person opens into is determined by their role, not by what they happen to discover in the settings.
A first-time user gets a chat box. A developer gets the workflow designer, MCP tooling and logs. Nobody has to grow into an interface they did not ask for.
Technical roles and teams for engineering, analysis roles for data work, executive personas for oversight — the library filtered to what is relevant.
Approved endpoints only. Sensitive roles can be held to models running inside your own network, while others reach hosted providers.
Group membership sets the base. Legal material is not reachable from an engineering role, and a contractor sees the project they are on and nothing else.
Ceilings and alerts per role, so heavy analytical work gets the headroom it needs and routine work does not quietly run up a bill.
Audit logging on the roles that need to be answerable, with the activity trail attached to the work rather than kept in a separate system.
A record of activity — which models were used, on what, and when.
Ceilings per role, team or provider, with warnings before the limit rather than after it.
The endpoint list is a decision, not a free-for-all. Users choose from what has been sanctioned.
Keys stored encrypted and injected automatically. Nobody pastes one into a chat or commits one to a repo.
Dev, staging and production kept distinct, so testing can never touch a live configuration.
How the app is actually being used, measured against the rules you have set for it.
At two people this is how you keep one client's documents out of another client's answers. At two hundred it's how you answer an auditor. Same mechanism either way.
You assemble the roles, models and memory that suit what you actually do.
Coding, review and architecture roles with your codebase in memory, and a local model carrying the volume.
Benchmark models against your real workload, then keep the analysis roles and datasets together in one project.
Schema documentation and standards in memory, on models you can run locally when the data cannot travel.
A project per client, memory kept separate, and a cost view telling you what an engagement actually consumed.
Enterprise mode for audit and budgets, with sensitive work pinned to models inside your own network.
Business mode, a team of three, and a template library that turns your best prompt into your default one.
Most of this is already happening, just spread across five tools and a folder of scripts.
| Instead of | You get |
|---|---|
| An LLM workbench for comparing models | The same comparison, wired into the app you actually work in |
| A separate coding assistant | Technical roles with your codebase already in memory |
| Prompt scripts scattered across notebooks and repos | A versioned template library with usage tracking |
| Pasting the same documents into chat every week | Memory that already has them |
| A subscription per person, per tool | One app, your own keys, your own models |
| No idea what any of it costs | Spend, tokens and response times per provider |
The workbench comparison is the closest one, and the difference is simple: a workbench is somewhere you go to test models and then leave. Here the testing sits inside the app you already work in, so the model you settled on is the one your roles are already running on.
ayoo.ai is in development. Join the waitlist and we'll let you know when builds go out — no drip campaign, no reselling your address.
macOS · Windows · Linux · Bring your own models