メインコンテンツへスキップ
0.682 PCC · 人間の合意を超える電話レベルのスコア

Brainiall Pronunciation
人間の合意を超える電話レベルのスコア

モデルがトレーニングされた人間のラッターよりも一貫している唯一の商業的な発音API。 0.590 生産レベルの光と、 0.682 AWS と GCP は同等の製品を持っていません; Azure は正確な番号を公開せずに 33 つの場所を提供しています。

Interactive Phoneme Scoring Demo

Click on individual phonemes below to inspect acoustic scores, ARPAbet alignment, and diagnostic feedback generated at runtime.

Interactive Phoneme Inspector
Sentence Score: 90 / 100
96
She
84
sells
91
seashells
98
by
76
the
93
seashore
Active Phoneme
/ʃ/
ARPAbet: SH
Acoustic Score
98/100
PCC Confidence: 0.682 (Exceeds 0.555 Human Agreement)
Diagnostic Feedback
Native-like articulation. Formant frequencies align with canonical reference.
Context word: “She

なぜ、この点数が重要なのか。

Traditional speech apps rely on vague sentence-level grades or subjective human reviews. Human raters only achieve a 0.555 Pearson Correlation Coefficient (PCC) with each other. Brainiall delivers objective, phone-level scoring that exceeds human agreement ceilings.

電話PCC(ライト)
0.590
+3.5pp 上の人間 (0.555) 9 モデル セット、CPU で実行します。
PCC(プレミアム)
0.682
+12.7pp 上の人間、 +2.5pp 上 SOTA (HIA 0.657)。
PCCの判決
0.711
+3.6pp 以上の人間 (0.675 ) 判決レベル XGBoost を介して電話+単語機能を上回る。
ラテンシー P95
423ms
シェアGPU上のライトレベル; プレミアム ~530ms 暖かい. 両方同期 — バッチは待たない。

二線統合

原始オーディオ(base64)を送信し、予想されるテキストを追加します 電話、単語、単語のスコアを返します - 1 HTTP 通話、ストリーミングは必要ありません。

POST /v1/pronunciation/assessJSON / Base64 Audio
curl -X POST https://api.brainiall.com/v1/pronunciation/assess \
  -H "Authorization: Bearer $BRAINIALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio": "<base64-wav-or-mp3>",
    "text": "She sells seashells by the seashore",
    "tier": "premium",
    "locale": "en-US"
  }'

# Response (deterministic score breakdown):
# {
#   "sentenceScore": 86,
#   "accuracyScore": 88,
#   "fluencyScore": 84,
#   "completenessScore": 100,
#   "wordScores": [
#     { "word": "she", "score": 96, "phonemes": [{ "phone": "SH", "score": 98 }, { "phone": "IY", "score": 94 }] },
#     { "word": "sells", "score": 84, "phonemes": [{ "phone": "S", "score": 92 }, { "phone": "EH", "score": 88 }, { "phone": "L", "score": 81 }, { "phone": "Z", "score": 75 }] }
#   ],
#   "latency_ms": 382
# }

When NOT to Use Pronunciation Assessment

  • Open-ended unstructured conversations without reference text: Pronunciation scoring requires the target text to compare acoustic phonemes against canonical alignment. For freeform transcription, use our Streaming STT engine in the Audio family.
  • Clinical motor speech disorder diagnosis: While acoustic scores provide objective metrics, clinical neurological evaluations require a licensed speech-language pathologist.
  • Tonal pitch or musical singing grading: Our scoring models are optimized for natural spoken linguistic phonemes, not musical melody or operatic pitch scales.

方法論

ライトレベルは9モデル集合(4×MLP+4×XGBoost+1× PhoneTransformer)で、Brainall Speech(エッジレベル)の音響モデルから抽出された機能を超えています - 私たちがSTT製品に使用する同じ背骨です。 ベンチマーク方法論 テストセット:テストセット 私たちの公共テストセット (英語 の L2 の SOTA の参照)

価格

ライトレベルの評価はCPU推論による本番品質で、プレミアムレベルは、私たちの公開テストセットで最高の公開精度を得るために、微調整された Brainiall Speech (Premium) エンジンを使用します。

フリー

ドル 0 / モー

100分/月 · ライトレベル · 永遠に無料

スタート

19ドル / 月

1000分/月 · ライトレベル · Word + フレーズ スコア

プロ

・99ドル/月

10,000分/月 · ライト + プレミアム · 電話レベルのスコア · 99.5% SLA

ビジネス

平成29年/月

50,000分/月 · 専用容量 · 創業者への直線

人間より一貫してスコアされた発音

Integrate phoneme-accurate speech scoring in minutes with instant API key issuance and zero monthly minimum.

無料トライアルを開始
Brainiall Pronunciation — 電話 PCC 0.59 / 0.68 (人間を超える) · Brainiall | Brainiall