Use vocal-emotion labels as a timeline, not a verdict.
diarize=true and num_speakers=2 for a two-person call. It returns timestamped speaker turns plus per-segment emotion and speaking-style labels. Oruk’s public MCP documents a no-account temporary test key limited to 30 minutes and three requests—no subscription required for that smoke test. Do not use those labels to infer anyone’s true feelings or who is “the problem.”[1][9][8]Why Oruk is the direct fit
- One analysis request can return a transcript, timestamps, speaker tags, 15 emotion labels and 16 speaking-style labels. Use the individual segments for a timeline; the recording-level summary is not a time series.
- Scores are independent labels, not shares of a single 100% total. A segment can score highly on more than one label. “Angry”, “frustrated”, “impatient” or “irritated” is a model description of vocal expression—not evidence of intent, a diagnosis, or someone’s private state.[1][3]
- Diarization only assigns temporary IDs such as
speaker_0; it does not identify people. Its docs warn about overlap, interruptions, short backchannels, and occasional wrong speaker attribution. Correct a few excerpts manually before counting anything. - Recorded Resonance is documented for English transcription; do not assume the English batch result generalizes to Turkish audio. The separate multilingual live route is preview and has its own limits.[1]
Cost and a reversible way to try it
The standard Hobby subscription is US$9/month for 250 speech-understanding minutes. The standard seven-day trial saves a payment method and charges the selected plan when it ends unless canceled in time. A short temporary-key test exists through Oruk’s public MCP tool, with no account required; use only a non-sensitive sample for the first smoke test. I did not create an account, start a trial, or send any audio.[2][9]
If audio is Turkish—or must stay on-device
emotion2vec+ supports utterance-level output and frame-level output at 50 Hz, with categories such as angry, neutral, sad and others. Its model card documents a local Python route through FunASR/ModelScope; I have not tested that route on macOS here. This keeps audio on your machine, but you must supply diarization, align scores to speaker turns, and create the timeline yourself. The model card does not establish accuracy on Turkish speech, these speakers, or natural arguments. Treat it as an experiment, not a proven judge. Check current model terms before redistributing or embedding it in another product.[5]
Hume’s Expression Measurement Tagger advertises batch analysis with 600+ expression and voice dimensions, but its product page asks users to contact the team for API access. It does not establish suitable language coverage, speaker diarization, or timestamp details for this use case. Confirm those, plus pricing and retention, before considering it. Hume positions its separate Prosody API for real-time emotion expression in live conversations.[7]
How to test the impression fairly
- Use a 10–15 minute representative excerpt with both speakers. First confirm language and audio quality. Preserve the original; run labels only on a copy.
- Review a timestamped transcript and speaker turns. Correct diarization and obvious transcription errors before interpretation.
- For each turn, inspect expression labels alongside what was actually said. Separately label concrete verbal behaviours: personal insult/devaluation, global accusation, contempt or sarcasm, interruption, defensive counterattack, clear boundary, apology/repair, and neutral disagreement.
- Apply the same rubric to A and B, ideally with names hidden during first scoring. Compare rates per 1,000 spoken words, not raw turn counts. Keep direct quotes and context beside any scored event.
- Inspect where voice and words disagree: angry-sounding prosody can accompany a legitimate boundary; calm delivery can still contain a personal attack. Emotion classifiers cannot determine which statement is hurtful or justified.
A 2025 review calls speech-emotion recognition one of the most difficult speech tasks: emotion definitions are broad, expression varies over time and within a dialogue, and estimates are inherently ambiguous. Individual differences may also warrant personalization. Treat model output as an exploratory listening index, not a reliable referee.[6]
Hume says its prosody scores are the model’s confidence that a speaker expresses a label through tone and language; it also warns that word-level scores depend heavily on context and are more stable at sentence level. That is still an inference about expression, not proof of a person’s private emotional state.[8]
Bottom line
For English audio and explicit speaker consent: use a short Oruk test, then manually audit the turns. For Turkish or a strict no-upload requirement: prototype the local diarization + emotion2vec route instead; expect engineering and language validation, not a polished turnkey answer. In neither case can a model “prove” bitterness or adversarial character. It can surface where to listen more carefully and support a symmetric review of both participants.