Skip to content
Projects

The World Doesn't Argue Fair

Side Project
Primary linkOpen live apphttps://bipinrimal314.github.io/dawym/
The World Doesn't Argue Fair
Consumer speech coaching apps are a graveyard. Speeko is dead ($315K total funding, iOS-only, gone). Poised got acqui-hired by Deepgram and rebranded as something nobody uses. Orai hasn't shipped an update since September 2024. Ummo counts filler words and does nothing else. LikeSo raised $0 and may or may not still exist. The survivors pivoted away from you. Yoodli raised $60M, hit $300M valuation, and sells exclusively to enterprises: Google, Snowflake, Databricks. BoldVoice found a defensible niche in accent reduction for non-native speakers and stopped there. Vocal Image coaches daily voice confidence from Estonia. What nobody survived building: a consumer speech analysis tool that tells you what you actually said, how you said it, and why it didn't land. Not a karaoke game with a score. Not a meeting assistant that counts your "ums." A forensic breakdown of your communication delivered without diplomatic hedging. So I built one. It runs in your browser. It doesn't need a server. It doesn't need your email.
You speak into your mic for anywhere from thirty seconds to ten minutes. DAWYM records, transcribes locally using OpenAI's Whisper (tiny, base, or small; you pick the accuracy-speed tradeoff), and returns a multi-layer analysis. Voice layer: Pitch variation over time (YIN algorithm, ~20Hz sampling). Volume dynamics in decibels. Pace mapped as words-per-minute in 15-second windows, color-coded by range. Pause detection that classifies silence as strategic or filler. Per-word clarity scoring that flags words you mumbled, rushed, or said too quietly; it tells you which words didn't land and why. Language layer: Filler words with timestamps and context (not just "you said 'um' 14 times" but where each one sits relative to your argument's structure). Hedging language detection: "I think maybe," "sort of," "kind of." Crutch phrases: "at the end of the day," "to be honest," "basically." Sentence length distribution. Readability scoring. Vocabulary density. The transcript itself has two views. Language Analysis highlights fillers, hedges, and crutches inline. "What Was Heard" colors every word by its clarity verdict: clear, soft, rushed, or mumbled. Wavy underlines on the unclear ones. Hover for the reason. Click any word to jump to that moment in the recording. The current word highlights during playback.
The browser SpeechRecognition API would have been easy. It's also wrong for this use case. Chrome-only. Sends your audio to Google's servers. Drops filler words from the transcript (which is the opposite of what you want when filler detection is the point). No word timestamps. No Firefox support. So DAWYM runs Whisper locally via Hugging Face's transformers.js. The tiny model is 40MB, downloaded once, cached forever. It transcribes in your browser with no server round-trip. The tradeoff: ONNX quantized models don't export cross-attention outputs, which means word-level timestamps aren't available. I approximate them by distributing time proportionally by character length within each segment. Good enough for filler detection and playback sync; not millisecond-precise. Pitch detection uses the YIN algorithm via pitchfinder. Volume is RMS calculation on time-domain data from the Web Audio API. Pace is arithmetic on the transcript plus timestamps. None of these are library calls you can npm-install and forget. The pitch detector needed careful tuning: smoothingTimeConstant at 0.3 (not the default 0.8, which dampens measured variation and lies to you about your range). autoGainControl disabled, because letting the browser dynamically adjust mic levels undermines every volume-dependent metric. The whole thing builds to 648KB of JavaScript plus 867KB of code-split transformers.js, plus the ONNX runtime. Twenty-seven source files, zero type errors, zero server dependencies.
Before your first real session, DAWYM captures two baselines. Step one: read a tongue-twister-loaded passage naturally. This establishes your baseline articulation, pace, and clarity. Step two: read a different passage as dramatically as possible; exaggerate everything. This captures your vocal range ceiling: the widest pitch variation, loudest volume, most dynamic pace you're capable of producing. Every session after that shows deltas against your baseline. Your filler count dropped 30% from your natural reading? Green arrow. Your pitch variation is narrower than your dramatic reading by 60%? That's data, not judgment; you decide what to do with it.
With a Claude API key (bring your own), DAWYM sends your full metrics and transcript to Haiku and gets back a 2-3 paragraph assessment that synthesizes everything the numbers show into language. The tone is not gentle:
"You spoke for 3:12. You were interesting for 1:40 of it. Your opening was a throat-clear; you didn't say anything worth hearing until 0:34. Your strongest sentence was at 2:15 but you buried it in a subordinate clause. You used 'basically' 6 times."
The name comes from watching a friend interact with his mother. She was stubborn, irrational, impossible to convince. He argued with her anyway, every time. Not to win; to practice articulating what he meant under pressure from someone who would never concede the point. I started noticing the same pattern everywhere. The people who were the sharpest communicators, the ones who could hold a room or dismantle an argument in real time, had all spent years doing exactly this: confronting authority figures, asking unwelcome questions, rebutting positions that weren't going to change. They weren't argumentative by nature. They were argumentative by practice. And that practice made them better at expressing ideas than anyone who only ever spoke to agreeable audiences. Socrates understood this. There's a story, possibly apocryphal, that he married Xanthippe precisely because she was the most difficult woman in Athens. When someone asked him why, he said: "If I can learn to live with her, I can learn to get along with anyone." The man who invented Western dialogue as a method of inquiry chose to go home every night to someone who would never, ever let him win an argument. That wasn't masochism. It was training. No presentation course teaches this. No speech coaching app models it. Every communication tool assumes you'll be speaking to a polite, receptive audience. That's recital, not communication. The world doesn't argue fair, and the only preparation is arguing with someone who won't. Phase 2 will add AI personas that argue back: The Mother, The Bureaucrat, The Boss. Each one unreasonable in a different way. Each one training a different skill.
The common belief is that the hard part of a speech analysis tool is the speech analysis. Pitch detection, filler counting, transcription: that's the interesting engineering. That belief is exactly backwards. The hard part is trust. Every metric you show to a user is a claim about their voice. If the filler detector flags "I like pizza" as a filler word, the user stops trusting the tool. If the clarity scorer reports words as "mumbled" because of approximated timestamps, the user questions every other metric. I spent more time on false positive reduction (50+ exemption rules for "like" alone, context-aware parsing for "just," "right," "actually") than on any single analysis module. A speech tool that's wrong 10% of the time isn't 90% useful. It's useless, because the user can never be sure which 10% to ignore.
Phase 1 complete and running. Phase 2 (Franklin Engine for vocabulary elevation, structural/rhetorical feedback, AI persona sparring) is next. Free tier stays free forever: no account, no server, no tracking. The entire local analysis engine is the acquisition channel. Try it: bipinrimal314.github.io/dawym Source: GitHub Stack: React 19, Vite 7, TypeScript, Tailwind v4, Whisper via transformers.js, pitchfinder, Recharts.