The market for AI transcription assistants changed character in 2026. The transcript itself has become a commodity, and value is migrating in two directions at once: downwards into the platform vendors' bundles, and upwards into the context layer for AI agents. In between sits a standalone segment that has to justify itself all over again.
This report does not sort the category by features. Features here have a half-life of a few months, because bot infrastructure and speech recognition can simply be bought in. Instead it sorts along two attributes that are architectural decisions rather than roadmap items: how a vendor gets to the audio, and who signs the cheque. That yields six strategic groups which barely compete with each other, even though most market overviews put them in the same table.
Disclosure: this report is published on Sally's blog, and Sally belongs to one of the six groups described. Wherever our own position is at stake we say so explicitly, including where the argument runs against us. The forecast on the sovereignty premium at the end of chapter five is one of those places.
Depending on why you are here, jump straight in:
- Strategy and investment: the six strategic groups and capital and consolidation
- Procurement and budget: pricing, bundling and what a seat really costs
- IT, data protection and works council: bot versus bot-less and codetermination as a critical path
- Project planning: why companies buy and how long it takes
- Concrete selection: the seven questions that decide a tender
The market in numbers
One caveat up front, and it qualifies everything else in this chapter: market figures in this category are extrapolations carrying substantial definitional uncertainty. They are useful for orders of magnitude and worthless as point forecasts.
Two estimates that differ by a factor of five
Two commonly cited data sets describe the same object and arrive at completely different results. The difference is not one of arithmetic quality but of scope: the narrow segment of pure note-taking and transcription tools, or the broad AI meeting assistant market including analytics and platform features.
| Metric | Narrow definition (meeting notes) | Broad definition (AI meeting assistants) |
|---|---|---|
| Market size 2025 | 623.5m USD | around 3.5bn USD |
| Market size 2026 | above 740m USD | around 4.31bn USD |
| Projection 2035 | 3.48bn USD | 34.28bn USD |
| Annual growth | 18.75 percent | 25.62 percent |
Anyone quoting one of these figures should quote the definition with it. A second effect undermines both projections: when Microsoft folds the feature into the bundle, revenue disappears from the market definition even as usage rises. That is happening right now, see chapter five. In this category a shrinking reported market can mean growing adoption.
Adoption falls as company size rises
The most striking figure in the adoption data is not the average but the gradient. Among solo professionals and small teams usage sits at 78 to 81 percent, while companies above 5,000 employees reach only 43 percent. Mid-market organisations between 1,001 and 5,000 employees land at 61 percent. Overall, 75 percent of professionals report using AI note-takers in work meetings, and 67 percent of Fortune 500 companies have a solution deployed somewhere in the organisation. These figures come from the benchmark report published by Laxis.
The gradient is the real finding: this category is adopted bottom-up, not rolled out top-down. Wherever procurement, data protection and codetermination apply, the organisation slows things down, not the benefit. That matches what 73 percent of companies name as the main barrier: privacy, not price or quality.
What usage actually delivers
On the benefit side the numbers are robust enough to plan with. 62 percent of users save roughly four hours a week, and the average professional spends around 146 hours a year reconstructing meeting context without AI support. Sales teams with full CRM sync report 8 to 12 hours saved per week and 15 to 20 percent of selling time recovered.
On accuracy it pays to keep two numbers apart. Leading models reach 95 to 97 percent word accuracy on clean audio, but only 85 to 90 percent with multiple speakers, crosstalk and accents. Speaker separation is reliable up to about eight distinct voices. An accuracy figure quoted without the recording situation is a marketing figure. For in-person meetings Sally states up to 98.8 percent when several phones act as distributed microphones in the room, and that is precisely the point: the recording situation decides, not the model.
The most honest maturity indicator comes from Microsoft
Taken together, the adoption figures paint an optimistic picture. A single data point corrects it. Microsoft reports around 15 million paid Copilot seats against a commercial base of more than 450 million Microsoft 365 users, which is under four percent penetration.
This is the vendor with the strongest distribution advantage in the world, pre-installed in practically every relevant company. If they sit below four percent, the entire category is still in its bottom-up phase rather than its enterprise rollout phase. For assessing growth forecasts, that is the single most important number in this report.
Six strategic groups instead of a feature table
A strategic group is a set of vendors that have made the same fundamental choices: similar architecture, similar buyer, similar pricing model. The practical use of the split is this: inside a group vendors substitute for each other, between groups they compete only indirectly. Putting Granola and Gong in the same comparison table compares two products that never fight over the same budget.
The two axes that cannot be copied
The first axis is capture architecture. Does the vendor enter the call as a visible bot participant, run as a local client on the machine, or work platform-natively from inside Teams, Meet and Zoom? This determines the sales argument, the compliance story and the cost structure, and it cannot be switched overnight.
The second axis is who signs the cheque: the individual with a credit card, the team lead with a small budget, IT and compliance in a large enterprise, or the chief revenue officer. Contract size, cycle length and what even counts as a buying argument all hang off this. A product built for the data protection officer rarely wins the credit card buyer, and the reverse holds too.
The six groups at a glance
| Group | Architecture | Buyer | Wins through |
|---|---|---|---|
| 1 Suite incumbents | platform-native | the contract already in place | distribution, no new vendor to onboard |
| 2 Horizontal bot platforms | bot in the call | individual, then team | integration breadth, free entry point |
| 3 Bot-less clients | local client | individual, then team | zero friction, no visible participant |
| 4 Revenue intelligence | bot in the call | chief revenue officer | pipeline analysis, coaching, forecasting |
| 5 Compliance-first regional | hybrid, bot and local | privacy, security, works council | processing agreement, EU hosting, transparency |
| 6 Vertical and suppliers | infrastructure | other vendors | bot infrastructure, speech recognition, domain depth |
Group 1 is Microsoft Copilot, Google Gemini in Meet and Zoom AI Companion. They need no bot because they are the platform. Their marginal price for one more summary is effectively zero, because the feature rides along in the bundle. They win not on product quality but on the fact that no new vendor has to be onboarded. This group sets the price anchor for everyone else without competing on price itself.
Group 2 is the horizontal bot platforms: Fireflies, Otter, tl;dv, Avoma and Read AI. Bot-based, grown product-led, strong on integration breadth. Fireflies has reached 16 to 20 million users depending on the source and a billion-dollar valuation, Otter has been in the market since 2016 with the brand recognition to match. In 2026 this group is in the most uncomfortable position, squeezed from below by free tiers and from above by bundling.
Group 3 is the bot-less, local-first clients, among them Granola, Fathom, Bluedot and Krisp. A local client captures the machine's system audio and nobody visibly joins the call. That solves several practical problems at once: no admitting the bot from the waiting room, no ten to thirty second join delay, and in environments where IT blocks bots at admin level it works at all.
Group 4 is revenue intelligence with Gong, Clari and Salesloft. For them the meeting is only a data source, the product is pipeline analysis and coaching. They sell to the revenue leadership rather than to IT, and the annual contract sits an order of magnitude above a note-taker. Important for market assessment: this group is relatively immune to Copilot bundling, because its value does not sit in the note. Around 27.7 percent of field sales teams already use conversation intelligence.
Group 5 is the sovereignty and compliance-driven regional vendors, in Germany including Sally and tucan.ai. The selling point is not the better summary but the defensible data processing agreement, EU hosting, transparency across the whole sub-processor chain and a written exclusion of AI training on customer data. Buying here is risk avoidance, not enthusiasm about productivity, and the buyer is the data protection officer or IT security. What Sally commits to on that front is set out on the page about GDPR and security.
One caveat belongs with this taxonomy: the groups are defined by architecture and buying motive, not by feature scope. A vendor can be bought as a group 5 product and still ship capabilities that have nothing to do with compliance. Sally is our own example of that: the buying motive is usually group 5, but the product reaches into the knowledge layer discussed in block 4 of the pricing chapter, with a searchable conversation archive and assistant features on top of it. The reverse holds too: treating group 5 as "privacy and nothing else" underestimates the group.
Why the supplier group explains the whole category
Group 6 is usually left out of market overviews because it sells no end-customer product. It is still the most important group in this report. It covers the vertical specialists for healthcare, legal and recruiting, but above all the infrastructure underneath: Recall.ai as bot infrastructure, plus AssemblyAI, Deepgram and ElevenLabs as the speech recognition layer.
The conclusion is uncomfortable and it carries the rest of this report. A startup can assemble a functionally competitive product for a low four-figure monthly spend. There are almost no technical barriers to entry in this category. It follows necessarily that the moat cannot architecturally sit in transcription. Anyone investing here and looking for the advantage in transcript quality is looking in the wrong place. The advantage sits in distribution, in the compliance position, in the accumulated data, or in embedding into existing processes.
Bot or no bot: the debate of the year
The dominant market narrative in 2026 is unambiguous: the bot has become a liability. The class action against Otter, now consolidated as In re Otter.AI Privacy Litigation, crystallised a suspicion buyers already held. A bot that joins and records by default is not a convenience feature, it is a risk. That is the best explanation for why the breakout product of the year uses no bot at all.
The number that finished the bot
The decisive finding is behavioural rather than legal. 84 percent of users change their behaviour or hold information back when they notice an AI bot in the call. Anyone who has watched a conversation shift the moment somebody says "I'll have this recorded" knows the effect first hand.
This is a conversational climate problem and a real one: a meeting in which the relevant information never gets said is worthless even with perfect transcription. For that problem, bot-less is the right answer.
Bot-less solves a social problem, not a legal one
This is where the common narrative parts company with the legal position. The obligations stay identical whether the recording sits visibly in the participant list or runs invisibly on a laptop. You still need a lawful basis under Art. 6 GDPR, participants still have to be informed under Art. 13, and in companies with a works council the enforceable codetermination right under § 87 (1) no. 6 of the German Works Constitution Act applies to any technical system capable of monitoring behaviour or performance.
Within the German legal framework, bot-less can even make things harder. An invisible recording satisfies the transparency obligation in a way that is more difficult to evidence than a join notification, because the notification is a documented signal while the absence of a bot is not. Anyone deploying bot-less has to organise transparency elsewhere, for instance through an explicit statement at the start and documented consent. What applies concretely in Germany, from § 201 of the Criminal Code to the role of the works council, is covered in our legal guide to recording meetings in Germany.
The governance gap in pure client tools
The second counterargument comes from IT departments. Many bot-less tools are by design personal capture apps with no organisational oversight: no central policy, no enforceable retention rules, no admin view of who recorded what. For one person on one device that works beautifully. The gap opens the moment a second person needs access or an admin wants to set a rule.
To be fair, that is not an argument against bot-less, it is an argument against bot-less without an administration layer. Which is exactly what this group is currently building.
Convergence points at hybrid architectures
From this follows the central forecast of this chapter. The market is not converging on bot-less as the winner but on hybrid architectures with a single admin policy across both capture modes. The movement is already visible: vendors now position themselves explicitly on offering both modes under the same admin controls, retention rules and security settings.
In practice: the capture mode becomes a setting the administrator picks per meeting type. Internal standups run locally, customer calls run with a visible notice, works council meetings are not recorded at all. Anyone who can only do one mode loses in enterprise deals. Sally covers both sides, with a bot in Google Meet, Zoom, Microsoft Teams and Webex and with the app for in-person meetings where no call exists in the first place.
Why and how fast companies buy
The decisive question is not what benefit the product delivers. It is this: what happens inside the company on the day somebody decides to spend money? Those are two different things, and the difference explains most lost deals.
Path A: shadow IT and the legalisation decision
The dominant case. One employee downloads a free tool, it works, they show colleagues, and within a few months fifty unauthorised instances are running in the company. Then IT security finds out. Formalisation typically follows about eighteen months after the start, usually triggered by someone in a security role noticing how widespread the practice has become.
The important part: in these cases the formal purchase is not a productivity decision but a legalisation decision. The company buys in order to bring existing, uncontrolled usage into a controlled frame. That also explains why the tool employees are already using often does not win the formal contract: once compliance runs the process, the selection criteria change completely.
Path B: the compliance trigger
An audit, a works council enquiry, an issue with the supervisory authority, or a professional confidentiality problem under § 203 of the German Criminal Code in law firms, medical practices and tax consultancies. This is the path that gets group 5 vendors into the tender at all. Without this trigger, the cheaper or the already present tool almost always wins.
Path C: the revenue trigger
Ramp time for new sales hires, CRM data quality, forecast hygiene, handovers between sales and delivery. Willingness to pay is highest here, because the benefit maps directly onto a revenue number. It is also why group 4 can work with contract values that would be unthinkable for a note-taker.
Path D: the knowledge trigger
Attrition and knowledge loss, distributed teams, multilingual work, integration after an acquisition. This path is gaining weight because it is the bridge to the agent story that explains the pricing logic in chapter five. Buyers on this path are not buying a transcript, they are buying a searchable archive.
The two brakes worth naming honestly
The first brake is codetermination, not price. Without a works agreement the rollout is contestable, the works council can enforce a halt under § 23 (3) of the Works Constitution Act, and in practice that means the results become unusable in employment matters. A defensible AI works agreement covers scope, a conclusively defined purpose, a closed list of approved tools, data protection including processing agreements, human oversight for personnel decisions, the training obligation under Art. 4 of the EU AI Act (mandatory since 2 February 2025), a complaints procedure, and term and adjustment. With cooperative counterparts, negotiation typically runs four to eight weeks.
The second brake is the bundling answer: "but we already have Copilot." Factually it is often wrong, because Copilot's meeting features depend on licence tiers and Teams usage and do not cover heterogeneous meeting platforms. As a conversation stopper it works anyway and delays deals by months. This dynamic decides the viability of the entire standalone segment.
How long a rollout realistically takes
An average across all segments would be worthless here. Only segmentation helps:
| Segment | Realistic cycle | Critical path |
|---|---|---|
| Individual, prosumer | days | none, self-service by credit card |
| SMB under 50 employees | 2 to 6 weeks | budget sign-off by management |
| Mid-market with works council | 3 to 9 months | works agreement |
| Large enterprise | 6 to 18 months | InfoSec review plus Copilot comparison |
One pattern is routinely underestimated in growth forecasts: pilot size does not scale linearly. A pilot with twenty users does not simply become a rollout with two hundred, because the pilot is carried by a business unit while the rollout triggers a fresh approval loop through IT and the works council. The second loop is often longer than the first. Anyone extrapolating growth from pilot conversion is systematically too optimistic.
Pricing: spread, not collapse
The naive answer on pricing is "prices are falling". The accurate answer is sharper: the category is splitting into a layer whose price trends to zero and a layer where prices rise. Five building blocks explain why.
Block 1: the base layer deflates
Transcript plus summary has become a commodity. Fathom runs an unlimited free tier with unlimited recording and storage, and Copilot Chat is included at no extra cost for every user on a qualifying Microsoft 365 plan. At the same time the marginal cost of a meeting hour has fallen into cents through cheaper speech recognition and inference. A 25 US dollar seat price for pure transcription is no longer defensible on a cost basis, and buyers notice.
Block 2: bundling is the structural break of the year
This is the most important development of 2026. On 1 July 2026 Microsoft moved Copilot out of the add-on model and into the core suite, making it official at Build 2026: Copilot is no longer a time-limited extra but a permanent part of the Microsoft 365 product line. The separate add-on price for smaller companies disappears into bundled plans, and several business and enterprise plans take price increases. At the top end the E7 Frontier Suite has been generally available since 1 May 2026 at 99 US dollars per user per month, including Agent 365 as a central control plane for discovering and securing AI agents across the organisation.
Why this is decisive for pricing: within a few quarters, the answer to "does the note-taker cost extra" will be no for Microsoft customers. Every standalone vendor then argues not against a low price but against a perceived price of zero. That is an entirely different sales situation and the harshest headwind for groups 2 and 3.
The arithmetic still rewards a second look, because that zero is a perceived one. Copilot costs 30 US dollars per user per month and cannot be bought on its own, it requires a qualifying base licence. With Microsoft 365 E3 at 39 US dollars the seat lands at 69, with E5 at 60 dollars it lands at 90. A 1,000 seat rollout on E3 therefore costs roughly 828,000 US dollars a year. On top comes an effect procurement regularly underestimates: 30 to 40 percent of licences sit unused after 90 days because seats are never assigned or never adopted. For comparison, Sally's own pricing starts at 8 euros per month.
Block 3: from seats to consumption
The second structural break concerns the billing model itself. Microsoft is leading: the Dynamics 365 pricing pages now show Copilot credits instead of fixed fees, and role-based AI is shifting into a consumption model. Copilot Studio for external use is billed through capacity packs at 200 US dollars for 25,000 credits.
The logic is straightforward: once agents are involved, cost correlates with usage rather than with headcount. Expect the whole category to move to a platform or base fee plus a consumption component for retrieval, agent runs and automations. For investors that matters because it changes revenue predictability and makes expansion revenue arise differently than through pure seat growth.
Block 4: value migrates upwards, away from the transcript
Here sits the actual investment case, and the people involved have stated it themselves. Granola's argument is that the value lies not in the notes themselves but in making the knowledge locked inside conversations accessible to other systems. If an AI agent can pull context from every meeting a team has ever had, it makes better decisions about the next step. Because note-taking becomes table stakes, the value lies in enabling actions out of notes and transcripts: drafting the follow-up, finding the slot, pulling knowledge from database and CRM to close a lead.
Translated: selling transcripts means selling a good with a falling price. Selling the searchable archive of all conversations as a context source means selling a good with a rising price, because value grows with the archive and switching gets expensive. Practically, that is the axis a knowledge base across all conversations sits on: not the individual transcript, but the searchable archive you can question in a chat instead of reading it. Sally covers exactly this layer and additionally makes the archive queryable by MCP straight from ChatGPT, Claude or Copilot. The point is not the interface but the direction of travel: context goes to where the work already happens instead of waiting in yet another tab. The 8,000+ tool connections and 99+ languages are the precondition for that, not the purpose.
Block 5: the sovereignty premium and its vulnerability
With compliance-driven buyers a premium of roughly 15 to 40 percent remains enforceable, because what is being bought is risk reduction rather than productivity, and risk reduction is considerably less price elastic.
Honesty requires the caveat, even though it argues against our own model: this premium is a function of how completely Microsoft and the US vendors close their EU data residency and sub-processor story. The more watertight that becomes, the smaller the uplift a regional vendor can charge. Anyone buying into group 5 today should therefore ask not only about the server location but about the whole sub-processor chain and about a written exclusion of AI training on their own data. Those are the parts that are harder to replicate than a data centre in Frankfurt.
Capital and consolidation
Funding now separates the groups more sharply than any feature set.
Who has fresh capital and who has not raised since 2021
On one side stands Granola: 125 million US dollars Series C at a 1.5 billion valuation in March 2026, led by Index Ventures and Kleiner Perkins, 192 million raised in total in under two years. In May 2025 the valuation still stood at 250 million. A sixfold valuation increase in ten months is not paid for notes, it is paid for the thesis in block 4. It is backed by 250 percent revenue growth in the quarter before the announcement and weekly retention above 70 percent.
On the other side sits the bot platform group. Otter has raised no primary capital since its Series B in February 2021 and reduced headcount to 278 by April 2026. Fireflies likewise has not raised primary capital since 2021 and stood at 117 employees in April 2026. Transaction data from YipitData adds that Granola is actively displacing Fathom, Otter and Fireflies across more than 900 mid-market companies, with threefold spend growth over six months, close to zero churn and a ninefold increase in gross adds from January to February 2026.
Three forecasts through 2027
First: list prices for standalone note-takers stay nominally stable because nobody cuts voluntarily, but effectively realised prices fall sharply through enterprise discounting. Measuring market health by list prices measures the wrong thing.
Second: revenue growth per customer over the next two years comes exclusively from add-on modules, not from base price increases. The base layer no longer permits an increase.
Third: anyone not established in the agent and governance layer by 2027 becomes an asset deal. A product that only delivers transcripts has no argument left against a bundle with a perceived price of zero.
Choosing in 2026: seven questions that decide
If feature lists are the weakest selection criterion, a tender needs different questions. These seven address exactly the things that cannot be rebuilt in a quarter:
- Both capture modes under one policy? Bot and local, with the same admin controls, retention rules and permissions.
- Where does the data sit, and who are the sub-processors? The server location alone is not an answer, the full chain is.
- Is the exclusion of AI training on your data confirmed in writing? Verbal assurances survive no audit.
- Does it cover your meeting reality? Meaning every platform you use plus in-person meetings, not just the one included in the bundle.
- Will it carry the works agreement? Purpose limitation, roles and permissions, deletion periods and human oversight have to be expressible in the product, not just in the contract.
- Does the archive come back out? API, export, connections to CRM and task systems. This decides whether in five years you hold a context asset or an archive.
- What is the all-in cost? Including base licences, consumption components and a realistic assumption about unused seats.
For the concrete tool comparison with prices and features there is a dedicated overview of the best AI meeting assistants compared. This report answers the question before that one: which group you should be shopping in at all.
Declaring our own interest, and measured with the same taxonomy: Sally sits in group 5 by buying motive and additionally on the knowledge layer by product. Concretely that means recording in Google Meet, Zoom, Microsoft Teams and Webex as well as through the app for in-person meetings, transcription in 99+ languages, up to 98.8 percent accuracy in in-person settings with several phones according to Sally, a searchable archive of every conversation with chat on top of it, access to that archive from AI assistants, connections to 8,000+ tools, and hosting exclusively in Germany by Aliru GmbH.
What Sally is not deserves the same clarity. Pipeline analysis, deal scoring and sales coaching are group 4 territory, and anyone buying on a pure revenue motive (path C) is better served there. Buyers on path A, B or D, who need to legalise grown shadow IT, resolve a compliance issue, or protect knowledge against attrition and distributed teams, should have Sally on the list. Sally can be tested free for 30 days, most meaningfully with your own real meetings rather than a demo.
Sources
Market figures, adoption rates and usage findings come from the benchmark report "The State of Meeting Note-Taking 2026" by Laxis and from the market projections by Market Research Future. Vendor categorisation and standalone pricing additionally draw on the market overviews by AI Tool Directory, Notta and Fellow. Those sources are published by market competitors and are cited here by name without a link.
- Granola Series C, valuation and market thesis: The Next Web on the 125 million round at a 1.5 billion valuation and Entrepreneur Loop with growth and retention figures
- Transaction data on market share and displacement: YipitData on Granola versus Fathom, Otter and Fireflies
- Microsoft pricing, bundling and licence prerequisites: Coworker on the 2026 enterprise pricing structure and Copilot Experts on tiers, the E7 Frontier Suite and unused licences
- Codetermination and AI works agreements: Skill Sprinters on § 87 BetrVG, mandatory components and the EU AI Act
- Legal position in Germany: Recording meetings: what is allowed in Germany
Report as of August 2026. Prices, valuations and licence terms change quickly in this category, particularly on the platform vendor side. For decisions with financial or legal weight, verify the figures with the respective vendor. This report is neither legal nor investment advice.




