Use vocal-emotion labels as a timeline, not a verdict.
diarize=true and num_speakers=2 for a two-person call. It returns timestamped speaker turns plus per-segment emotion and speaking-style labels. Oruk’s public MCP documents a no-account temporary test key limited to 30 minutes and three requests—no subscription required for that smoke test. Do not use those labels to infer anyone’s true feelings or who is “the problem.”[1][9][8]Why Oruk is the direct fit
- One analysis request can return a transcript, timestamps, speaker tags, 15 emotion labels and 16 speaking-style labels. Use the individual segments for a timeline; the recording-level summary is not a time series.
- Scores are independent labels, not shares of a single 100% total. A segment can score highly on more than one label. “Angry”, “frustrated”, “impatient” or “irritated” is a model description of vocal expression—not evidence of intent, a diagnosis, or someone’s private state.[1][3]
- Diarization only assigns temporary IDs such as
speaker_0; it does not identify people. Its docs warn about overlap, interruptions, short backchannels, and occasional wrong speaker attribution. Correct a few excerpts manually before counting anything. - Recorded Resonance is documented for English transcription; do not assume the English batch result generalizes to Turkish audio. The separate multilingual live route is preview and has its own limits.[1]
Cost and a reversible way to try it
The standard Hobby subscription is US$9/month for 250 speech-understanding minutes. The standard seven-day trial saves a payment method and charges the selected plan when it ends unless canceled in time. A short temporary-key test exists through Oruk’s public MCP tool, with no account required; use only a non-sensitive sample for the first smoke test. I did not create an account, start a trial, or send any audio.[2][9]
If audio is Turkish—or must stay on-device
emotion2vec+ supports utterance-level output and frame-level output at 50 Hz, with categories such as angry, neutral, sad and others. Its model card documents a local Python route through FunASR/ModelScope; I have not tested that route on macOS here. This keeps audio on your machine, but you must supply diarization, align scores to speaker turns, and create the timeline yourself. The model card does not establish accuracy on Turkish speech, these speakers, or natural arguments. Treat it as an experiment, not a proven judge. Check current model terms before redistributing or embedding it in another product.[5]
Hume’s Expression Measurement Tagger advertises batch analysis with 600+ expression and voice dimensions, but its product page asks users to contact the team for API access. It does not establish suitable language coverage, speaker diarization, or timestamp details for this use case. Confirm those, plus pricing and retention, before considering it. Hume positions its separate Prosody API for real-time emotion expression in live conversations.[7]
Local runs on M1 / 16 GB; existing cloud tools cover different layers
Follow-up, 5 October 2026: For a private first experiment, local classification is feasible. That changes the practical recommendation: do not add another cloud account just to learn whether a local vocal-expression timeline is possible. A complete two-speaker workflow still needs transcription, speaker-turn attribution, and manual review.
Its Sentiment Analysis returns positive / neutral / negative per sentence, with start/end times, confidence, and speaker labels when diarization is enabled. The documentation explicitly says it interprets transcript text. It is useful for a wording-sentiment timeline but cannot establish angry delivery, bitterness, or sarcasm from vocal acoustics. Use sentiment_analysis=true and speaker_labels=true for that limited purpose; verify language and model compatibility separately.[10]
Scribe v2 provides word timestamps, diarization, and audio-event tags such as laughter. I found no documented batch emotion-probability timeline in the checked STT response contract. Expressive mode does use emotional/prosodic cues for live agent turn-taking, but that is not an exported offline emotion-classifier result. Its emotional generation tags direct synthetic voice output; do not confuse them with detection. Zero-retention STT requests are enterprise-only.[11][16][17]
Gemini accepts recordings and supports timecode/speaker transcription with audioTimestamp=true. A practical experiment is to supply short timestamped clips and ask separately for observable delivery and wording, preserving uncertainty. That is prompted qualitative analysis—not a documented calibrated per-emotion probability endpoint. Gemini Live's enable_affective_dialog changes how an interactive model responds to tone; it is not a ready-made offline score export. No cloud audio request or billing was initiated for this research.[12][13]
Tested: superb/wav2vec2-base-superb-er, pinned to revision 441a7599c3b22107314dcbd9166621c5c83f2cc5, using existing PyTorch/Transformers. No new dependency packages or Hermes configuration changes. Public demo clips only; no private audio. Hardware probe: Apple M1, 8 CPU cores, 16 GB RAM; MPS available.
- 94.6 million parameters; English IEMOCAP baseline with four categories: neutral, happy, angry, sad.[14]
- Two clips of 5.13 s and 4.39 s, three runs per clip/device. Offline-model rerun: warmed CPU inference 0.123–0.141 s; warmed Apple GPU/MPS inference 0.072–0.077 s.
- Peak process resident memory reported by macOS
ru_maxrss: 0.96 GiB on the offline rerun; initial download/load plus inference process peaked at 1.21 GiB. These are process measurements, not a promise about the complete pipeline or GPU memory accounting. - Model and samples downloaded once, then model loading repeated with Hugging Face/Transformers offline modes. No runtime cloud inference service was used.
- Accuracy is the limitation, not this short-clip compute test: a demo advertised as neutral scored happy 0.544 versus neutral 0.452. These are the classifier's softmax scores, not calibrated certainty. No independent natural-conversation or language validation was performed.
Emotion inference alone works on this machine. A complete private two-speaker timeline still needs speech segmentation/diarization, optional local transcription, and alignment. Do not run an hour-long waveform as one transformer input; process short, speaker-homogeneous clips sequentially.
Next local model to evaluate: emotion2vec+ base, not large, is documented at approximately 90 million parameters and includes nine output categories. Its frame mode extracts features at 50 Hz; those are not automatically 50 calibrated emotion decisions per second. Hardware fit is plausible, but its FunASR route and end-to-end memory/speed have not been tested here. Do not transfer the SUPERB timings to it.[15]
Working recommendation: use the verified local baseline to test the timeline machinery privately; use AssemblyAI only for wording/transcript context if cloud processing is acceptable. Vertex is a possible qualitative second reviewer with actual audio, not a numerical ground truth. Neither the hosted alternatives nor the local baseline has been validated on the user's recordings.
Follow-up sources
How to test the impression fairly
- Use a 10–15 minute representative excerpt with both speakers. First confirm language and audio quality. Preserve the original; run labels only on a copy.
- Review a timestamped transcript and speaker turns. Correct diarization and obvious transcription errors before interpretation.
- For each turn, inspect expression labels alongside what was actually said. Separately label concrete verbal behaviours: personal insult/devaluation, global accusation, contempt or sarcasm, interruption, defensive counterattack, clear boundary, apology/repair, and neutral disagreement.
- Apply the same rubric to A and B, ideally with names hidden during first scoring. Compare rates per 1,000 spoken words, not raw turn counts. Keep direct quotes and context beside any scored event.
- Inspect where voice and words disagree: angry-sounding prosody can accompany a legitimate boundary; calm delivery can still contain a personal attack. Emotion classifiers cannot determine which statement is hurtful or justified.
A 2025 review calls speech-emotion recognition one of the most difficult speech tasks: emotion definitions are broad, expression varies over time and within a dialogue, and estimates are inherently ambiguous. Individual differences may also warrant personalization. Treat model output as an exploratory listening index, not a reliable referee.[6]
Hume says its prosody scores are the model’s confidence that a speaker expresses a label through tone and language; it also warns that word-level scores depend heavily on context and are more stable at sentence level. That is still an inference about expression, not proof of a person’s private emotional state.[8]
Bottom line
For a private first experiment, local short-clip emotion inference is now verified on M1 / 16 GB. Use the baseline to build and audit a timeline before adding another hosted tool. Oruk remains a direct hosted English option if speaker consent and retention terms are acceptable. AssemblyAI covers transcript sentiment; Vertex is a possible qualitative audio reviewer; ElevenLabs supplies transcript/speaker/event context. The local emotion2vec+ base candidate is not yet benchmarked. No model here can prove bitterness or adversarial character.