Recorded speech · emotion over time · two-speaker analysis

Use vocal-emotion labels as a timeline, not a verdict.

Research checked 5 October 2026. This is a general tool evaluation; no private conversation or recording is included or analyzed.

Best practical first test—if the recording is in English: Oruk Resonance file analysis with diarize=true and num_speakers=2 for a two-person call. It returns timestamped speaker turns plus per-segment emotion and speaking-style labels. Oruk’s public MCP documents a no-account temporary test key limited to 30 minutes and three requests—no subscription required for that smoke test. Do not use those labels to infer anyone’s true feelings or who is “the problem.”[1][9][8]
Privacy gate: Oruk’s terms require the rights and speaker consent needed to submit a recording. With diarization enabled, an external processor may retain uploaded audio for up to 48 hours; Oruk says it has no SOC 2, ISO 27001, independent third-party security audit, or published DPA for the standard offering. For intimate recordings, do not upload until both speakers agree and those terms are acceptable.[3][4]

Why Oruk is the direct fit

English recordings · batch API · speaker turns

Cost and a reversible way to try it

The standard Hobby subscription is US$9/month for 250 speech-understanding minutes. The standard seven-day trial saves a payment method and charges the selected plan when it ends unless canceled in time. A short temporary-key test exists through Oruk’s public MCP tool, with no account required; use only a non-sensitive sample for the first smoke test. I did not create an account, start a trial, or send any audio.[2][9]

If audio is Turkish—or must stay on-device

Local building block · not a ready-made conversation app

emotion2vec+ supports utterance-level output and frame-level output at 50 Hz, with categories such as angry, neutral, sad and others. Its model card documents a local Python route through FunASR/ModelScope; I have not tested that route on macOS here. This keeps audio on your machine, but you must supply diarization, align scores to speaker turns, and create the timeline yourself. The model card does not establish accuracy on Turkish speech, these speakers, or natural arguments. Treat it as an experiment, not a proven judge. Check current model terms before redistributing or embedding it in another product.[5]

Hosted batch alternative · access is not self-serve

Hume’s Expression Measurement Tagger advertises batch analysis with 600+ expression and voice dimensions, but its product page asks users to contact the team for API access. It does not establish suitable language coverage, speaker diarization, or timestamp details for this use case. Confirm those, plus pricing and retention, before considering it. Hume positions its separate Prosody API for real-time emotion expression in live conversations.[7]

How to test the impression fairly

  1. Use a 10–15 minute representative excerpt with both speakers. First confirm language and audio quality. Preserve the original; run labels only on a copy.
  2. Review a timestamped transcript and speaker turns. Correct diarization and obvious transcription errors before interpretation.
  3. For each turn, inspect expression labels alongside what was actually said. Separately label concrete verbal behaviours: personal insult/devaluation, global accusation, contempt or sarcasm, interruption, defensive counterattack, clear boundary, apology/repair, and neutral disagreement.
  4. Apply the same rubric to A and B, ideally with names hidden during first scoring. Compare rates per 1,000 spoken words, not raw turn counts. Keep direct quotes and context beside any scored event.
  5. Inspect where voice and words disagree: angry-sounding prosody can accompany a legitimate boundary; calm delivery can still contain a personal attack. Emotion classifiers cannot determine which statement is hurtful or justified.

A 2025 review calls speech-emotion recognition one of the most difficult speech tasks: emotion definitions are broad, expression varies over time and within a dialogue, and estimates are inherently ambiguous. Individual differences may also warrant personalization. Treat model output as an exploratory listening index, not a reliable referee.[6]

Hume says its prosody scores are the model’s confidence that a speaker expresses a label through tone and language; it also warns that word-level scores depend heavily on context and are more stable at sentence level. That is still an inference about expression, not proof of a person’s private emotional state.[8]

Bottom line

For English audio and explicit speaker consent: use a short Oruk test, then manually audit the turns. For Turkish or a strict no-upload requirement: prototype the local diarization + emotion2vec route instead; expect engineering and language validation, not a polished turnkey answer. In neither case can a model “prove” bitterness or adversarial character. It can surface where to listen more carefully and support a symmetric review of both participants.