Skip to content
Writing

What Three ML Models See When They Watch You Speak

March 16, 20263 min read
Technical
DAWYM analyzes what you say and how you say it. The Teleprompter helps you deliver a script. But every speaking coach I've ever worked with spends half their evaluation on something neither tool measures: what your body is doing. "Stop swaying." "Use your hands." "Look at us, not the floor." "You're fidgeting with your ring again." Vanessa Van Edwards' research at Science of People found that viewers who watched only the first seven seconds of TED Talks rated speakers nearly identically to those who watched the full talk. Your body speaks before your mouth does. The audience decides if they trust you before you finish your opening sentence. No consumer tool measures this properly. PowerPoint's Speaker Coach checks if you're visible on camera. Yoodli does voice analysis. Orai counts filler words. The closest thing is uSpeek, which evaluates 11 body language criteria from your webcam, but it's early-stage and upper-body only. So I built the thing that watches every second, never forgets a fidget, and tracks progress over weeks.
Three MediaPipe models run simultaneously in a Web Worker on an OffscreenCanvas: Pose Landmarker (33 body keypoints, ~6MB). Shoulders, elbows, wrists, hips, knees. This is how the app knows if you're swaying, slouching, or standing with your weight on one foot. Face Landmarker (478 landmarks + 52 blendshapes, ~5MB). Tracks eye direction (camera or notes?), facial expression (engaged or frozen?), and micro-expressions. Hand Landmarker (21 landmarks per hand, ~5MB). Distinguishes between an open illustrative gesture and fidgeting with your watch. Total: ~18MB, cached after first download. All inference runs in the browser via a Web Worker, keeping the main thread free for UI. On Apple Silicon, 15-25 combined frames per second. No audio or video leaves your device.
Synthesized from Toastmasters Pathways evaluation criteria, TED coaching methodology, Nick Morgan's Power Cues, and Van Edwards' research:
  1. Eye Contact. 50-70% audience-facing is good. Below 50% reads as disengaged.
  2. Gesture Quality. The "Clinton Box" (shoulder-to-shoulder, chest-to-waist) is the optimal gesture zone.
  3. Movement. Purposeful repositioning vs. nervous swaying. Tracked via center-of-mass over time.
  4. Posture. Shoulder alignment, spinal angle, weight distribution.
  5. Expression. Facial engagement via blendshape data. Smiling, frowning, raising eyebrows, or flat affect.
  6. Stillness. The anti-fidget score. Self-touching, repetitive hand movements, weight shifting.
Level 1: Foundation. Eliminate distractors. Stop swaying. Stop fidgeting. Keep your hands visible. Look at the camera. The app flags every occurrence and shows you the count. Level 2: Intentional Basics. Maintain posture for a full minute. Use gestures in the Clinton Box. Hold eye contact for 3-5 second intervals. Level 3: Integration. Coordinate gesture with speech rhythm. Move purposefully to mark transitions. Vary facial expression. Level 4: Mastery. All six dimensions flowing simultaneously. Body language stops being a checklist and starts being presence.
Body language norms vary. Americans expect strong eye contact; in many East Asian cultures, sustained direct gaze toward authority figures is disrespectful. A body language coach that hard-codes American norms and applies them globally is a culturally specific tool pretending to be universal. Stage Simulator has adjustable cultural profiles that shift scoring thresholds for eye contact, gesture amplitude, movement range, and personal space. You train for the audience you'll actually face.
While building this, I realized the measurement engine doesn't care what skill it's scoring. MediaPipe tracks body position. The scoring rubric decides what "good" means. For a speaker, good posture means shoulders back and weight centered. For a dancer, it means something entirely different. Same 33 keypoints, same movement tracking, different definition of "correct." The same app that coaches you through a presentation could coach you through a salsa basic or a bharatanatyam stance. The architecture is already there.
Stage Simulator is the third piece of a speaking suite that now covers voice (DAWYM), script delivery (Teleprompter), and body language (Stage Simulator). Three apps, zero overlap. Phase 6 integrates all three into combined sessions. Stack: React 19, Vite, TypeScript (strict), Tailwind v4, MediaPipe Pose/Face/Hand Landmarker, Web Workers, OffscreenCanvas. Code: Stage Simulator on GitHub