Building speech interfaces to restore access_
Our optical speech interface turns the sound of a room into readable language on the wearer's own lens. Speech is captured by a beamformed array in the frame, transcribed as it is spoken, and drawn into one eye by a waveguide — end of sentence to glass in under a second.
This technology is built to return conversational access to people who are deaf or hard of hearing, and it is built the only way a device worn among other people should be: with no camera, no recording and no idea who is speaking.
A room is only accessible if you can follow what is said in it.
Captioning stops at the edge of a screen.
Hearing aids amplify. They do not disambiguate.
A phone held up between two faces is not a conversation.
Access should not depend on everyone else agreeing to change how they speak.
What Snow cannot do.
Everything below is a property of the instrument rather than a preference inside an app. A frame with no image sensor cannot be made to see by a later software update, and a pipeline that never forms a voiceprint has none to leak. We design from the position of the person who did not buy the device and is simply in the room.
-
No camera.
There is no image sensor anywhere in the frame. Nobody around the wearer is ever seen, framed, photographed or filmed, and no part of the system takes visual input of any kind.
-
No recording.
Audio exists as a few hundred milliseconds of buffer that is continuously overwritten. A segment becomes text and the audio is gone. There is no file, no tape, no transcript history and nothing to export, hand over or subpoena.
-
No identification.
Snow does not know who is speaking and is not built to find out. No voiceprint, speaker profile or biometric template is computed, stored or matched — not for the wearer, and not for anyone else in the room.
-
No inference about people.
Snow reports what was said. It does not score, rank, rate or draw conclusions about the person who said it: no emotion recognition, no sentiment, no intent, no truthfulness, no profiling. The lens carries language and nothing else.
-
No speakers.
Nothing is played back into the room. What the wearer reads is legible only to the wearer, and is not visible from outside the lens.
-
No wake word.
The frame is not waiting to hear its own name, which is the only way to be certain it is not waiting to hear yours.
-
No assistant.
Snow answers no questions and volunteers nothing. It is an instrument the wearer reads, not a participant in the conversation.
-
No accounts, no history.
There is no profile to build and no conversation log to accumulate. What was said in a room stays in that room, because the system keeps no record that it happened.
Snow is designed against the GDPR and the EU AI Act from the first drawing rather than audited against them afterwards. The categories that attract the most scrutiny in a wearable — biometric identification, emotion recognition, profiling, retention — are not mitigated here. They are absent, and the architecture on this page is what makes them absent.
One for hearing the room. One for understanding it.
Same language in, same language out.
Speech in the room becomes text on the lens, verbatim, in the language it was spoken. This is the shorter path through the instrument and the one the accessibility case rests on.
- Fidelity
- Verbatim
- Stages used
- Array · Onset · Loki · Glass
- Works offline
- Yes
Their language in, yours out.
For a conversation held in a language the wearer does not read. One additional stage, and the only one that ever sees text rather than sound.
- Fidelity
- Interpreted
- Stages used
- Array · Onset · Loki · Mimir · Glass
- First languages
- ĺslenska · Polski · English · Spanish · French · German · Italian · Portuguese · Russian · Turkish
End of a sentence to words on glass, in under a second.
Five stages. Mics in the brow. A gate that decides what is speech. A listener that turns it into text. An interpreter, if you need one. Glass that draws the line into your left eye.
45 grams
A wire rim.
Noise cancelling
Listen where you're facing.
Powerful AI
Automatic Speech Recognition.
Titanium finish
High strength metal
Realtime captions
Smart glass speech interface
A budget, not a benchmark. The live rig on this page measures the real thing and prints what it cost. Playback here runs at quarter speed so the stages are readable.
A frame you'd wear anyway.
Under 45 grams, because a device nobody will wear restores nothing. Every component below earns its mass: the array, the driver, the cell and the waveguide, and no sensor that is not required to read a sentence.
- Weight
- < 45 g
- Display & Optics
- Geometric Reflective Waveguide
- Indicator
- Red microLED
- Input
- 4-mic Beamforming array
- Prescription
- Optional
- Status
- Working prototype
Built with the people who study this.
Snow is a speech-technology and optics programme based in Reykjavík. The hard parts are not secret, and none of them are solved. We are looking for research partners — in speech recognition and low-resource languages, in applied machine learning, in optics and electrical engineering — and for deaf and hard-of-hearing people in Iceland to tell us where the instrument fails them.
Íslenska on a power budget.
The only listener that hears Icelandic is the slowest one we have, and a frame carries a cell rather than a rack. Closing that gap — a low-resource language, running at conversational latency, on a wearable power envelope — is the problem we would most like to work on with a PhD student.
- Fields
- ASR · low-resource speech · on-device inference
Under a second, honestly.
A caption that arrives late is a caption the wearer reads while the next sentence is being spoken. We publish a budget and measure against it, including partial hypotheses shown before a speaker has finished — and the open question of what a caption should do when the model is unsure.
- Fields
- streaming ASR · uncertainty · HCI
Type at the edge of vision.
A waveguide gives you a small, bright, fixed window. How much text a person can actually read inside it while looking at somebody's face — and what to drop when there is more language than window — is an optics question and a legibility question at the same time.
- Fields
- waveguide optics · legibility · accessibility research
Privacy that survives contact.
A device worn among people who did not consent to it has to be defensible in the room, not only in a policy. We want the no-camera, no-recording, no-identification architecture examined by people who will try to break the claim rather than repeat it.
- Fields
- privacy engineering · law · applied ethics
Tell us who you are, not just where to write.
Snow is a small team in Reykjavík. If you supervise students, run a lab, fund this kind of work, or would be the one wearing the frame, we would rather hear from you early than launch at you late. Direct: snow.is@outlook.com
We'll write when there is something real to look at — the first frames go out for testing in Iceland, with the people who need this most.