b-d.io / Notes
Live captions on smart glasses that don't flicker
Speech arrives word by word and changes its mind. The lens shouldn't jump every time it does.
What goes wrong
The naive approach — send the last few sentences as one block of text whenever anything changes — looks terrible on a lens:
- the firmware re-wraps the block, so lines shift sideways and up and down;
- the screen fills up and then a whole block jumps away at once;
- a sentence that gets shorter (an English line replaced by its shorter Chinese translation) pulls everything below it upward;
- a burst of recognition updates turns into a burst of Bluetooth packets, and the lens lags behind the speaker.
1. Do the layout on the phone
Break every line yourself, using the same glyph widths and line-breaking rules as the firmware, then send lines the glasses won't re-wrap. Leave a pixel or two of margin so a line you measured as fitting never wraps again on the glasses.
- Even G2 publishes its font metrics (in the
@evenrealities/pretextpackage): per-glyph advances in 1/16 px, kerning, and a fixed 27 px line height.
2. Treat the lens as a window on a stream of lines
Keep every caption wrapped into lines, newest at the end, and show a window of the last N lines. Then move the window down only, one line at a time:
- Settle at a resting number of lines (we use 6 on G2). Beyond it, move up one line, wait about a second, move again.
- If the newest line would otherwise fall off the bottom, move sooner (we use ~0.12 s) — but still one line per step.
- Keep one empty line at the bottom so the next line always has somewhere to appear.
- Never move the window back up. If lines above get shorter, leave the space at the bottom.
3. Reserve rows for each sentence
A sentence takes the most rows it ever needed. When it settles into fewer lines — or its translation is shorter — keep the extra rows as blank lines. Nothing below it moves.
4. Mark what's still being spoken
A lens usually draws everything at one brightness, so you can't dim the words still being recognized. We open every sentence with › and close the one still being spoken with …; the closing mark disappears when the sentence is final.
5. One update per tick, always the newest
Don't send on every recognition callback. Keep the current lens text in memory and run a loop that sends at most one update per tick, of whatever is newest. A backlog collapses into a single draw instead of replaying every intermediate version.
- Even G2: replace a text container's content in place; don't rebuild the page.
- Meta Ray-Ban Display: we send at most one update every 0.35 s.
6. Test it like a viewer, not a log
Keep a "shadow screen" on the phone that renders exactly the lines you send, at the lens's size. You'll see jumps and reflows immediately — and you can develop without wearing the glasses.