NEWS
42CAP Bets €7.7 Million on Deepslate’s European Voice Models
42CAP is backing Deepslate’s EU-hosted speech-to-speech model so insurers and contact centres can keep call audio in Europe.
Berlin lab Deepslate has raised €7.7 million to sell speech-to-speech voice AI that stays inside the EU. Munich firm 42CAP led the seed, with Alstin Capital, existing backer SIVentures and several business angels.
The 15-person team trains its own audio-to-audio model for European languages and already runs it for insurers, contact centres and platforms. Independent tests timed Deepslate Opal at 440 milliseconds to first audio, while the new money is aimed at German names, dialects and more EU compute, not at topping the global reasoning charts.
A €7.7 Million Seed From Munich and Berlin
Deepslate presents itself as a model lab, not a telephony app. It trains the model, then sells access through a self-serve API, volume deals and self-hosting. 42CAP general partner Julian von Fischer said Europe needs its own language models and that Deepslate is one of the few teams on the continent training a speech-to-speech system rather than wrapping someone else’s. He cited a small team and a fraction of the compute large labs burn.
Alstin Capital partner Andreas Schenk said co-founders Paskal Paesler and Jan Brachthäuser had already scaled latency-sensitive systems for millions of users, and that Alstin is backing a European voice lab it thinks can compete worldwide. The cheque is still seed money. It has to pay for training data, a sales bench the company says cannot keep up with demand, and more production kit in European data centres.
THE ROUND IN BRIEF
- The cheque: A €7.7 million seed led by 42CAP, with Alstin Capital, SIVentures and angels.
- The lab: Fifteen people in Berlin, ISO 27001 certified, models hosted in the EU.
- The product: Deepslate Opal, a native speech-to-speech model sold via API, volume plans and self-hosting.
- The spend: European training data, sales and marketing, and more capacity in EU data centres.
That mix is a wager on regulated call work. If German insurers and contact centres will not send audio to a US cloud, a fast local model with a price list in euro cents is the product they can sign. If they will, Deepslate is just another mid-pack name on a public board dominated by larger labs.
Three Seconds Was Too Slow for a Meeting Bot
Paesler and Brachthäuser started in 2023 with Unter Anderem, a meeting assistant that listened in and answered from company knowledge. The voice stack of the time took about three seconds to reply. That gap killed a natural turn. Spoken words went to text, a language model wrote a reply, and a third system read it out. Each hop added delay. Tone, irony and dialect dropped out in the middle.
They spent months on then-obscure open research and built the stack themselves. The result is a three-part system trained in-house: a speech encoder that maps the waveform into a semantic and prosodic space, a reasoning core based on an open-weights language model that Deepslate post-trains per language, and a speech decoder that speaks from those embeddings. Trainable projectors join the encoder and decoder to the language model, so a new core can be swapped without retraining the whole stack. The company says that swap takes days, not months.
FROM MEETING BOT TO MODEL LAB
- 2023: Paesler and Brachthäuser launch Unter Anderem; a three-second voice lag makes live talk unusable.
- January 2026: The company brings Deepslate Opal out of stealth as its first enterprise speech-to-speech model.
- June 23, 2026: Artificial Analysis lists Opal as the fastest model in its new speech-to-speech index, at 0.44 seconds to first audio.
- October 1, 2026: Deepslate closes the €7.7 million seed and posts a 2-cent launch price on its site.
US-trained models, the founders argue, still trip on German names, street addresses and dialects. In insurance and banking, a swapped house number is an expensive error. That is the unglamorous problem the architecture is supposed to attack, and it is also where the new training budget is going.
How Fast Is Deepslate Opal on Artificial Analysis?
On the Artificial Analysis speech-to-speech leaderboard, Deepslate Opal’s time to first audio is 0.44 seconds. That is the figure the company prints on its homepage, and it is still far quicker than OpenAI’s GPT-Realtime-2 High at 1.14 seconds and Google’s Gemini 3.8 Live Extended Thinking High at 1.35 seconds. Krafton’s Raon SpeechChat now sits at 0.04 seconds, so Opal is no longer the fastest row on the live board. It remains the speed name among models that also post a serious reasoning score.
In June, the same bench was blunter. Artificial Analysis wrote that Opal had the fastest average time to first audio in the index at 0.44 seconds.
Announcing the Artificial Analysis Speech to Speech Index, our new synthesis metric for native Speech to Speech model quality, comprising of Big Bench Audio, Full Duplex Bench, and 𝜏-Voice
The index provides a single measure of how well native Speech to Speech models perform,… pic.twitter.com/WW6yVwO19R
— Artificial Analysis (@ArtificialAnlys) June 23, 2026
Reasoning is the other column, and it is less kind. Opal scores 85% on Big Bench Audio. GPT-Realtime-2 High and Grok Voice Think Fast 2.0 High both sit at 97%. Gemini 3.8 Live Extended Thinking High is at 98%. On τ-Voice, which scores replica customer-service tasks, Opal’s agentic result is 17.5%. Grok’s high setting is 56.5%. Gemini’s extended-thinking live model is 68.6%. Conversational dynamics, which tracks pauses, interruptions and backchannels, put Opal at 85.7%, behind the OpenAI and Alibaba realtime lines.
SPEED, REASONING AND TASK SCORES
| Model | Time to first audio | Speech reasoning | Agentic score |
|---|---|---|---|
| Raon SpeechChat | 0.04 s | 58% | – |
| Deepslate Opal | 0.44 s | 85% | 17.5% |
| Gemini 2.5 Flash Native Audio Dialog | 0.63 s | 69% | – |
| Grok Voice Think Fast 2.0 High | 0.70 s | 97% | 56.5% |
| GPT-Realtime-2 High | 1.14 s | 97% | 39.8% |
| Gemini 3.8 Live Extended Thinking High | 1.35 s | 98% | 68.6% |
The company still calls Opal the fastest speech-to-speech model Artificial Analysis had measured as of September 2026, and von Fischer used the same “fastest in the world” line at close. The live table now has a quicker row in Raon, whose reasoning score is 58%. Deepslate’s own lab test, from the end of a user’s utterance to the first syllable and without long-haul network delay, is 250 milliseconds. Those are different clocks. The public one is how time to first audio is measured on Big Bench Audio.
The X argument around that board is a quality fight among Grok, OpenAI and Google. Deepslate shows up as the latency specialist. That split matches a product built for phone queues, and it leaves the reasoning crown to labs with far more compute.
Contact Centres Buy Hosting They Can Audit
The paying user is not a consumer chatting with a toy voice. Deepslate says insurers, contact centres and platforms already run the model in production. Its telecom page pitches native SIP integration for European carriers, G.711, and a stack that can sit on the customer’s network so call audio never leaves. Self-hosting can be air-gapped. In that mode, Deepslate says it has no access to the data.
Cloud copies run in Deutsche Telekom’s German cloud. The company is ISO 27001 certified and says it publishes where compute happens, who the subprocessors are, and what is still open. Paesler put the sales line in one breath.
For our customers, data sovereignty is not a nice-to-have, it is a prerequisite. And either it can be verified or it is worthless. That is why we disclose where the computing happens, who our subprocessors are and what is still open
Paskal Paesler, co-founder, Deepslate
Banks, clinics and public bodies sit in the same constraint. They often cannot send customer audio to a US provider, even one with a Frankfurt region, because jurisdiction is not the same as a rack location. Deepslate’s bet is that those buyers will take a slightly weaker agentic score if the audio path is inspectable. The risk is that a contact-centre vendor still wants the tool-calling rates Gemini and Grok post on τ-Voice, and will try to paper over residency with contracts instead of a local model.
The Data Programme for Names and Dialects
The round’s first technical job is not another parameter bump. Deepslate will expand model training and a European data programme, with a particular focus on German street names, personal names and dialects, while it tries to cut latency and lift voice quality. That is the work US English-heavy corpora skip, and it is the work that makes a claims line or a directory-assistance prompt usable.
The company also says Opal posted the best error rate in its comparison field for European languages on CoVoST 2. That figure is Deepslate’s own reading of the field, not a number Artificial Analysis publishes. The underlying CoVoST 2 multilingual speech corpus covers 21 languages into English and English into 15, released as a research set for speech translation. It is a reasonable exam for European coverage. It is not a live insurance call with a Bavarian accent and a hyphenated surname.
Sales capacity is the other bottleneck. Demand, the company says, already exceeds what the current team can work. Part of the €7.7 million hires sales and marketing. Another part scales production in European data centres as call volume grows. Long range, Deepslate talks about voice as the interface in vehicles, devices and robots, and says some customers already use the model as a speech front end for their own systems. The near-term bill is still names, dialects and minutes on the phone.
Pay-As-You-Go Minutes and a Model You Can Swap
The homepage now sells the audio-in, audio-out model with no middleware to any builder. A launch banner prices usage at 2 cents a minute, against a regular 10 cents, excluding VAT, and the company warns the offer can change for new sign-ups after the launch window. Enterprise deals are custom, aimed at contact centres, insurers and platforms running millions of minutes. Developers get official plugins for LiveKit and Pipecat, Python and Node SDKs, plus raw WebSocket, WebRTC and SIP, so one realtime model can replace a chained speech-to-text, language-model and text-to-speech setup.
Deepslate’s product copy puts a cascaded pipeline at about 800 milliseconds, with each conversion adding 200 to 300 milliseconds, and says text in the middle strips sarcasm and compounds errors on accents, names and addresses. Opal is meant to keep rhythm, emphasis and intonation because it never dumps the call to a transcript. Whether a Berlin lab can hold that quality while it swaps cores through projectors is the engineering half of 42CAP’s bet. The commercial half is whether 2-cent minutes, a Telekom rack and a German names corpus beat a better English reasoner that legal will not approve.
HOW BUYERS CAN RUN OPAL
- Self-serve API: Pay as you go at the posted 2-cent launch rate, 10 cents once the offer lapses, no contract call required.
- Volume and SLA: Custom enterprise terms for contact centres, insurers and platforms that need guaranteed capacity.
- Self-host: The full stack on the customer’s network, air-gapped if required, with no outbound path and no Deepslate access to audio.
The seed does not make Opal the smartest voice model in the world, and the live latency table no longer gives it a clean first place. It does pay for the unfashionable layer European phone queues actually fail on: street names, dialects, and a compute path a compliance team can read. That is the wager 42CAP just made.
Frequently Asked Questions
Does Deepslate Need Separate Speech-to-Text and Text-to-Speech?
No. Opal is a native speech-to-speech model, so builders do not bolt on a separate recogniser, language model or speaker. The company ships LiveKit and Pipecat plugins, Python and Node SDKs, and raw WebSocket, WebRTC and SIP, and says one realtime model replaces the three-stage chain.
What Does the €15 Starter Credit Buy on Deepslate?
The site’s launch FAQ says a €15 starter credit covers 750 minutes at 2 cents a minute, with 10 concurrent calls included in the release offer. That math is 15 divided by 0.02, and the company says the promotional rate may change for new sign-ups after the launch period.
Where Does Deepslate Process Call Audio?
The company says all data is processed and stored in the EU, with cloud copies in Deutsche Telekom’s German cloud and an option for customers to self-host, including with no external connection. It is ISO 27001 certified and says a data-processing agreement, subprocessor list and hosting region are open for review.
What Is CoVoST 2 in Deepslate’s Language Claim?
CoVoST 2 is a research speech-translation set with 21 languages into English and English into 15, published under a CC0 licence so labs can test multilingual speech systems. Deepslate says Opal had the best error rate in its comparison field for European languages on that set; it has not published a single word-error number beside that claim.
-
BUSINESS4 months agoMusk’s $914 Billion Lead Is a Public SpaceX Bet
-
NEWS2 months agoMicrosoft’s 96% Cyber Score Sits at 86.3% on the Board
-
SPORTS3 months agoFree Live Sports Streaming in 2026: What to Watch Without Cable
-
ENTERTAINMENT4 weeks agoSterling Point Holds No. 2 on Prime Video After 32 Days
-
NEWS4 weeks agoAn AI Lung Drug Shifted Biological Age Clocks in Patients
-
ENTERTAINMENT3 months agoThe Odyssey’s IMAX Film Run Hit a 41-Theater Limit
-
NEWS3 months agoPixel 10’s Amazon Cut Now Sits Beside Pixel 11
-
NEWS6 months agoWhite Sands Footprints Put People in Ice Age New Mexico
