Recorded speech · emotion over time · two-speaker analysis

Use vocal-emotion labels as a timeline, not a verdict.

Research checked 5 October 2026. No private recording analyzed. Follow-up: existing providers and a verified local M1 benchmark.

Best direct hosted test—if the recording is in English: Oruk Resonance file analysis with diarize=true and num_speakers=2 for a two-person call. It returns timestamped speaker turns plus per-segment emotion and speaking-style labels. Oruk’s public MCP documents a no-account temporary test key limited to 30 minutes and three requests—no subscription required for that smoke test. Do not use those labels to infer anyone’s true feelings or who is “the problem.”[1][9][8]
Privacy gate: Oruk’s terms require the rights and speaker consent needed to submit a recording. With diarization enabled, an external processor may retain uploaded audio for up to 48 hours; Oruk says it has no SOC 2, ISO 27001, independent third-party security audit, or published DPA for the standard offering. For intimate recordings, do not upload until both speakers agree and those terms are acceptable.[3][4]

Why Oruk is the direct fit

English recordings · batch API · speaker turns

Cost and a reversible way to try it

The standard Hobby subscription is US$9/month for 250 speech-understanding minutes. The standard seven-day trial saves a payment method and charges the selected plan when it ends unless canceled in time. A short temporary-key test exists through Oruk’s public MCP tool, with no account required; use only a non-sensitive sample for the first smoke test. I did not create an account, start a trial, or send any audio.[2][9]

If audio is Turkish—or must stay on-device

Local building block · not a ready-made conversation app

emotion2vec+ supports utterance-level output and frame-level output at 50 Hz, with categories such as angry, neutral, sad and others. Its model card documents a local Python route through FunASR/ModelScope; I have not tested that route on macOS here. This keeps audio on your machine, but you must supply diarization, align scores to speaker turns, and create the timeline yourself. The model card does not establish accuracy on Turkish speech, these speakers, or natural arguments. Treat it as an experiment, not a proven judge. Check current model terms before redistributing or embedding it in another product.[5]

Hosted batch alternative · access is not self-serve

Hume’s Expression Measurement Tagger advertises batch analysis with 600+ expression and voice dimensions, but its product page asks users to contact the team for API access. It does not establish suitable language coverage, speaker diarization, or timestamp details for this use case. Confirm those, plus pricing and retention, before considering it. Hume positions its separate Prosody API for real-time emotion expression in live conversations.[7]

Local runs on M1 / 16 GB; existing cloud tools cover different layers

Follow-up, 5 October 2026: For a private first experiment, local classification is feasible. That changes the practical recommendation: do not add another cloud account just to learn whether a local vocal-expression timeline is possible. A complete two-speaker workflow still needs transcription, speaker-turn attribution, and manual review.

AssemblyAI · text sentiment, not acoustic emotion

Its Sentiment Analysis returns positive / neutral / negative per sentence, with start/end times, confidence, and speaker labels when diarization is enabled. The documentation explicitly says it interprets transcript text. It is useful for a wording-sentiment timeline but cannot establish angry delivery, bitterness, or sarcasm from vocal acoustics. Use sentiment_analysis=true and speaker_labels=true for that limited purpose; verify language and model compatibility separately.[10]

ElevenLabs · transcript, speakers, non-speech events

Scribe v2 provides word timestamps, diarization, and audio-event tags such as laughter. I found no documented batch emotion-probability timeline in the checked STT response contract. Expressive mode does use emotional/prosodic cues for live agent turn-taking, but that is not an exported offline emotion-classifier result. Its emotional generation tags direct synthetic voice output; do not confuse them with detection. Zero-retention STT requests are enterprise-only.[11][16][17]

Vertex / Gemini · qualitative audio review

Gemini accepts recordings and supports timecode/speaker transcription with audioTimestamp=true. A practical experiment is to supply short timestamped clips and ask separately for observable delivery and wording, preserving uncertainty. That is prompted qualitative analysis—not a documented calibrated per-emotion probability endpoint. Gemini Live's enable_affective_dialog changes how an interactive model responds to tone; it is not a ready-made offline score export. No cloud audio request or billing was initiated for this research.[12][13]

Local · actually exercised, including offline model loading

Tested: superb/wav2vec2-base-superb-er, pinned to revision 441a7599c3b22107314dcbd9166621c5c83f2cc5, using existing PyTorch/Transformers. No new dependency packages or Hermes configuration changes. Public demo clips only; no private audio. Hardware probe: Apple M1, 8 CPU cores, 16 GB RAM; MPS available.

  • 94.6 million parameters; English IEMOCAP baseline with four categories: neutral, happy, angry, sad.[14]
  • Two clips of 5.13 s and 4.39 s, three runs per clip/device. Offline-model rerun: warmed CPU inference 0.123–0.141 s; warmed Apple GPU/MPS inference 0.072–0.077 s.
  • Peak process resident memory reported by macOS ru_maxrss: 0.96 GiB on the offline rerun; initial download/load plus inference process peaked at 1.21 GiB. These are process measurements, not a promise about the complete pipeline or GPU memory accounting.
  • Model and samples downloaded once, then model loading repeated with Hugging Face/Transformers offline modes. No runtime cloud inference service was used.
  • Accuracy is the limitation, not this short-clip compute test: a demo advertised as neutral scored happy 0.544 versus neutral 0.452. These are the classifier's softmax scores, not calibrated certainty. No independent natural-conversation or language validation was performed.

Emotion inference alone works on this machine. A complete private two-speaker timeline still needs speech segmentation/diarization, optional local transcription, and alignment. Do not run an hour-long waveform as one transformer input; process short, speaker-homogeneous clips sequentially.

Next local model to evaluate: emotion2vec+ base, not large, is documented at approximately 90 million parameters and includes nine output categories. Its frame mode extracts features at 50 Hz; those are not automatically 50 calibrated emotion decisions per second. Hardware fit is plausible, but its FunASR route and end-to-end memory/speed have not been tested here. Do not transfer the SUPERB timings to it.[15]

Working recommendation: use the verified local baseline to test the timeline machinery privately; use AssemblyAI only for wording/transcript context if cloud processing is acceptable. Vertex is a possible qualitative second reviewer with actual audio, not a numerical ground truth. Neither the hosted alternatives nor the local baseline has been validated on the user's recordings.

Follow-up sources

  1. AssemblyAI Sentiment Analysis
  2. ElevenLabs STT capabilities
  3. Google audio understanding
  4. Google Live affective-dialog configuration
  5. SUPERB emotion-classifier model card
  6. emotion2vec+ base model card
  7. ElevenLabs STT endpoint contract and retention
  8. ElevenLabs expressive live-agent mode

How to test the impression fairly

  1. Use a 10–15 minute representative excerpt with both speakers. First confirm language and audio quality. Preserve the original; run labels only on a copy.
  2. Review a timestamped transcript and speaker turns. Correct diarization and obvious transcription errors before interpretation.
  3. For each turn, inspect expression labels alongside what was actually said. Separately label concrete verbal behaviours: personal insult/devaluation, global accusation, contempt or sarcasm, interruption, defensive counterattack, clear boundary, apology/repair, and neutral disagreement.
  4. Apply the same rubric to A and B, ideally with names hidden during first scoring. Compare rates per 1,000 spoken words, not raw turn counts. Keep direct quotes and context beside any scored event.
  5. Inspect where voice and words disagree: angry-sounding prosody can accompany a legitimate boundary; calm delivery can still contain a personal attack. Emotion classifiers cannot determine which statement is hurtful or justified.

A 2025 review calls speech-emotion recognition one of the most difficult speech tasks: emotion definitions are broad, expression varies over time and within a dialogue, and estimates are inherently ambiguous. Individual differences may also warrant personalization. Treat model output as an exploratory listening index, not a reliable referee.[6]

Hume says its prosody scores are the model’s confidence that a speaker expresses a label through tone and language; it also warns that word-level scores depend heavily on context and are more stable at sentence level. That is still an inference about expression, not proof of a person’s private emotional state.[8]

Bottom line

For a private first experiment, local short-clip emotion inference is now verified on M1 / 16 GB. Use the baseline to build and audit a timeline before adding another hosted tool. Oruk remains a direct hosted English option if speaker consent and retention terms are acceptable. AssemblyAI covers transcript sentiment; Vertex is a possible qualitative audio reviewer; ElevenLabs supplies transcript/speaker/event context. The local emotion2vec+ base candidate is not yet benchmarked. No model here can prove bitterness or adversarial character.