A voice & speech AI research lab

Teaching machines to listen.

Speech is the oldest interface we have, and still the hardest one to build for. KogniGen studies how machines hear it — the acoustics, the structure hiding inside them, and what it takes to answer well enough to be worth talking to.

From pressure in the air to a voice that answers

01 — The core

It starts as pressure in the air.

Before it is language, a voice is a wave: a few thousand changes in air pressure every second, carrying an accent, a mood, a room, and a head cold all at once. Everything a speech model ever knows about you arrives in that one thin stream.

We work on the front of that pipe — capturing it cleanly, and holding on to what matters when the room is noisy and two people talk at once.

02 — Deep analysis

Sound becomes structure.

Split the wave by frequency over time and speech stops looking like noise. Vowels show up as stacked bands, consonants as bursts and hiss. This is where a model learns that ship and sheep differ by a few dozen milliseconds in one band.

Our work here is about representations that survive contact with reality: far-field microphones, crosstalk, compression, and speakers no training set ever met.

03 — The conversation

Structure becomes meaning.

Words are only half of it. When you pause, whether your pitch rises, how quickly you come back — that is where intent lives, and it is the part most systems throw away. A model that transcribes perfectly and interrupts constantly is still a bad listener.

So we study turn-taking as a first-class problem: when to speak, when to wait, and how to recover when it gets it wrong.

04 — Global resonance

Meaning becomes a voice.

Then it has to answer — out loud, in time, and in a voice that carries meaning rather than just pronouncing it. Emphasis in the wrong place is its own kind of error.

Most of the world's languages have almost no recorded data, and most speech systems quietly serve a handful of them well. We think that is the interesting problem, not an edge case to handle later.

What we research

Four problems we keep coming back to.

  • Recognition in the wild

    Benchmarks are quiet, close-miked and take turns politely. Kitchens, classrooms and cars are none of those things. We work on speech recognition that holds up in noise, at distance, across accents, and when people talk over each other.

  • Synthesis with intent

    Getting the phonemes right is table stakes. We study prosody — stress, rhythm, pitch, the pause before the important word — because that is what makes a synthetic voice understandable rather than merely intelligible.

  • Conversation in real time

    In dialogue, a late answer is a wrong answer. We treat latency, barge-in and turn-taking as research problems in their own right, not engineering details to optimise once the model works.

  • Languages without data

    Most languages have no transcribed corpus worth the name, and speakers of those languages are not a rounding error. We work on transfer, self-supervision and evaluation for low-resource speech.

How we work

Positions we don’t trade away.

  • A voice is identity

    A recording of someone speaking is biometric, not just audio. We keep as little as the work allows, we don’t sell or mine it, and cloning a voice requires the consent of the person it belongs to.

  • Evidence over demo

    Anyone can cut a good demo reel. We aim to publish methods, evaluation and the cases where our models fail — including the accents and conditions they handle worst.

  • Every accent is first-class

    A system that works beautifully for some speakers and poorly for others isn’t finished, it’s selective. We measure across speakers, not just on average.

  • Say what it can’t do

    Speech interfaces fail invisibly — they mishear and answer confidently. We’d rather a system admit uncertainty than perform fluency it hasn’t earned.

Where we are

Early, and saying so.

A lab that overstates its progress has already told you something about its standards. So here is the plain version.

No benchmark claims yet
We’re not publishing error rates or comparisons until there is a write-up behind them that you can check. Numbers without a method are just decoration.
Pre-launch
Bubbles is not on sale and there is no price to quote you. Write to us and we’ll tell you when that changes.
Hiring, quietly
If you work on speech — recognition, synthesis, or the awkward real-time parts in between — we’d like to hear from you before we write a job post.

Where the research goes

Someone has to talk to it.

Bubbles is our product: an embodied companion that children speak to out loud. It is also the hardest test we could have picked for the lab — children mumble, shout, interrupt, invent words, and wander off mid-sentence. Speech systems tuned on adult dictation fall apart in a playroom.

Everything we learn there feeds back into the research, and everything the research produces has to survive a five-year-old before we believe it.

Visit Bubbles

Get in touch

Tell us what we’re getting wrong.

Collaborate with the lab, join it, or send us the recording our models will hate most. One of us reads every message.

kognigen@gmail.com Meet Bubbles