A physically modeled throat that sings vowels, drawn as a midsagittal anatomy plate.
The airway of the head on screen is the instrument. Its outer wall (pharynx, soft palate, hard palate, teeth) is fixed and its front wall is a tongue body, root, blade, jaw and lips in the style of Mermelstein's articulatory model; 44 rays cast from the wall measure the width of each 0.4 cm section, which becomes an area function. That area function drives a Kelly-Lochbaum digital waveguide, pressure waves scattering at every change of area, stepped twice per audio sample, with a nasal side branch behind the velum joined by a three-way junction, and the source is the Liljencrants-Fant glottal flow model shaped by one voice quality parameter. The spectrum panel is the FFT of the impulse response of that same lattice, so the labeled formants are the ones you hear, and the vowel chart is filled in the background by measuring this tract's own vowel space. The airway glows with the RMS pressure in each section, the standing waves of the formants, so the demo reads even with the sound off.
Try it. Drag the tongue body, the lips (up and down to open, sideways to round) or the velum (down to open the nose). Drag in the vowel chart to jump to a vowel. Keys A, E, I, O and U pick vowels, N toggles the velum, arrows change the pitch, Space stops and starts the voice, M mutes; sound starts on the first click and the autopilot sings when you let go.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a vowel synthesizer that models the human vocal tract as a tube, in JavaScript with the HTML canvas element and the Web Audio API. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Model the tract as 44 short tube sections from the vocal folds to the lips, each with a cross-sectional area. Start from a uniform tube and add a "tongue": a smooth bump that narrows the tube, with a position and a height you can change.
- Implement a Kelly-Lochbaum waveguide: each section holds a right-going and a left-going pressure wave. Every step, at each joint, part of the wave reflects with k = (A1 - A2) / (A1 + A2) and the rest passes on. The lips reflect most of the wave back inverted (about -0.85) and the output is what leaks out; the glottis end reflects about +0.75. Run two steps per audio sample.
- Drive it with a glottal pulse train at about 120 Hz (a simple rising-then-falling pulse each period is fine at first) plus a little noise.
- Generate samples in chunks into AudioBuffers and schedule them back to back. Create the AudioContext only after a click, and show a "Click for sound" hint until then.
- Draw the area function as a tube profile, and let the mouse drag the tongue bump along the tube and up and down. You should hear the vowel change.
Once that works, make it beautiful:
- Compute the tract's frequency response by feeding an impulse through the same lattice and taking an FFT, and plot it with the formant peaks labeled F1, F2, F3.
- Add an F1 versus F2 chart with reference vowels (i, e, a, o, u) and a dot that moves as you drag.
- Make the tube glow with the RMS pressure in each section, so you can see the standing waves even with the sound off.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as the Liljencrants-Fant glottal model for a more natural voice, a nasal branch with a velum, or an autopilot that sings vowel melodies.