Real-time Transcription (Speech-to-Text)

    Stream audio and receive low-latency Speech-to-Text results over WebSocket. This endpoint is the recommended starting point for live transcription and automatically detects the spoken language.

    Recommended endpoint

    wss://api.valsea.ai/v1/realtime/asr

    Connect here first. Omit language or set it to auto for automatic detection.

    Automatic language detection is the default

    One unified API for every language and accent. Stream audio to a single endpoint and VALSEA automatically detects the spoken language—no language selection required.

    valsea-rtt automatically detects the spoken language by default, so you do not need to select an input language in advance — language can be omitted from the session.start message entirely. You can make this explicit by setting language to "auto".

    {
      "type": "session.start",
      "model": "valsea-rtt"
    }
    

    VALSEA supports 181 languages, dialects, and variants overall; coverage varies by endpoint. Auto-detect supports multilingual conversations and code-switching, including language changes within a conversation or sentence. If you already know the input language, you can pin it with a specific language code — see Use a fixed input language at the bottom of this page.

    Connection

    Endpoint: wss://api.valsea.ai/v1/realtime/asr

    Authentication

    You must authenticate the WebSocket connection by passing your API key in the HTTP headers during the handshake.

    Headers:

    • Authorization: Bearer YOUR_API_KEY (Recommended)
    • X-API-Key: YOUR_API_KEY (Supported)

    Paid Rate-Limit Bypass

    If you need to exceed the realtime connection RPM limit for a session, you can bypass the rate-limit check by sending one of these opt-in flags during the WebSocket handshake:

    • Header: X-Bypass-Rate-Limit: true
    • Query parameter: bypass_rate_limit=true
    • Query parameter: bypassRateLimit=true

    Bypass applies only to the per-organization rate-limit check. Authentication, credit checks, and session billing still apply. RTT sessions using this bypass are billed at 2x the normal realtime credit cost. The initial session.created event includes rateLimitBypass: true and billingMultiplier: 2 when bypass is active.

    Message Flow

    1. Connect: Client establishes WebSocket connection.
    2. Session Created: Server sends session.created event.
    3. Start Session: Client sends session.start to configure language and model.
    4. Stream Audio: Client sends audio.append messages with base64-encoded PCM16 audio chunks.
    5. Receive Transcripts: Server streams transcript.partial and transcript.final events.
    6. Commit Audio: Client sends audio.commit when a user stops speaking (optional/VAD dependent).
    7. Stop Session: Client sends session.stop to end the session.

    Partial vs Final (Important)

    RTT emits two transcript event types for each utterance:

    • transcript.partial: low-latency, in-progress text. This can change as more audio arrives.
    • transcript.final: stable text for a completed segment. Treat this as the committed result.

    Recommended client behavior:

    1. Keep a temporary currentPartial string for transcript.partial.
    2. Append only transcript.final to your persisted transcript history.
    3. Clear currentPartial when you receive a matching transcript.final.

    Client Messages

    session.start

    Initialize the session with configuration.

    {
      "type": "session.start",
      "model": "valsea-rtt",
      "hint_text": "Optional context or vocabulary",
      "enable_correction": true,
      "language": "auto",
      "language_hints": ["en", "vi", "ar"],
      "noise_suppression": "off"
    }
    
    FieldTypeDescription
    modelstringModel to use (e.g., valsea-rtt).
    hint_textstringOptional list of words or context to improve accuracy.
    enable_correctionbooleanEnable post-processing for grammar/language correction (default: true).
    languagestringOptional input language. Defaults to auto for code-switch-aware automatic detection.
    language_hintsstring[]Optional two- or three-letter codes that bias automatic detection without restricting it. Maximum: 20. Also accepted as languageHints.
    noise_suppressionstring"off" (default) or "rnnoise" — runs incoming audio through noise suppression before transcription.

    audio.append

    Send audio data.

    {
      "type": "audio.append",
      "audio": "BASE64_ENCODED_PCM16_DATA"
    }
    
    • Format: Raw PCM 16-bit, 16kHz (recommended), mono.
    • Encoding: Base64 string.

    You may also send raw binary PCM16 frames directly over the WebSocket. Binary frames are treated as audio.append messages by the server.

    audio.commit

    Signal the end of a speech segment (e.g., VAD triggered silence).

    {
      "type": "audio.commit"
    }
    

    session.stop

    End the session gracefully.

    {
      "type": "session.stop"
    }
    

    Server Messages

    session.created

    Sent immediately upon connection.

    {
      "type": "session.created",
      "sessionId": "rtt_...",
      "supportedModels": ["valsea-rtt"]
    }
    

    session.ready

    Sent when the backend engine is connected and ready to receive audio.

    {
      "type": "session.ready",
      "sessionId": "rtt_...",
      "engine": "valsea-7"
    }
    

    transcript.partial

    Intermediate transcription results (low latency, may change).

    {
      "type": "transcript.partial",
      "text": "The quick brown",
      "isFinal": false,
      "timestampMs": 1888
    }
    

    transcript.final

    Finalized text for a speech segment. On this endpoint, text and rawText are always identical — translation fields (translated, sourceLanguage, targetLanguage) never appear here.

    {
      "type": "transcript.final",
      "text": "The quick brown fox jumps over the lazy dog while the market opens at 9.30 in the morning.",
      "rawText": "The quick brown fox jumps over the lazy dog while the market opens at 9.30 in the morning.",
      "isFinal": true,
      "timestampMs": 6328
    }
    

    error

    Sent when an error occurs.

    {
      "type": "error",
      "code": "INVALID_MESSAGE",
      "message": "Failed to parse message"
    }
    

    Error codes on this endpoint

    CodeFires whenRecoverable?
    AUTH_REQUIREDNo API key provided on connectNo — reconnect with a key
    AUTH_FAILEDInvalid API key, or key has no associated organizationNo — reconnect with a valid key
    RATE_LIMITEDOrg exceeded RTT connections-per-minuteYes — retry after retryAfterMs
    INSUFFICIENT_CREDITSOrg credit balance is ≤ 0No — until credits are topped up
    INVALID_MESSAGEA client message couldn't be parsed as JSONYes — session stays open
    ALL_ENGINES_FAILEDNo transcription backend could be initialized for the requested languageNo — session ends
    NOT_READYaudio.append sent before session.readyYes — wait for session.ready
    AUTO_ROUTE_ERRORAuto-detection failed for a specific utterance (language: "auto" only)Yes — subsequent turns unaffected

    Event Handling Pattern

    Use this pattern to avoid duplicated or unstable transcript content:

    let currentPartial = '';
    const finalSegments = [];
    
    ws.on('message', (raw) => {
      const msg = JSON.parse(raw);
    
      if (msg.type === 'transcript.partial') {
        currentPartial = msg.text || '';
      }
    
      if (msg.type === 'transcript.final') {
        finalSegments.push(msg.text || '');
        currentPartial = '';
      }
    });
    

    Example (Node.js)

    const WebSocket = require('ws');
    const fs = require('fs');
    
    const ws = new WebSocket('wss://api.valsea.ai/v1/realtime/asr', {
      headers: { 'X-API-Key': 'YOUR_KEY' },
    });
    
    ws.on('open', () => {
      ws.send(
        JSON.stringify({
          type: 'session.start',
          model: 'valsea-rtt',
        }),
      );
    });
    
    ws.on('message', (data) => {
      const msg = JSON.parse(data);
    
      if (msg.type === 'session.ready') {
        const audioStream = fs.createReadStream('audio.raw');
        audioStream.on('data', (chunk) => {
          ws.send(
            JSON.stringify({
              type: 'audio.append',
              audio: chunk.toString('base64'),
            }),
          );
        });
      } else if (msg.type === 'transcript.final') {
        console.log('Final:', msg.text);
      }
    });
    

    Use a fixed input language

    By default, sessions auto-detect the input language. If you already know the language being spoken, set it explicitly for the most consistent routing:

    {
      "model": "valsea-rtt",
      "language": "singlish"
    }
    

    Language Hints

    Automatic detection is recommended. Omit language or send "language": "auto" in session.start. If you know the input language, you may send a supported code such as english, singlish, vietnamese, arabic, arabic-egypt, or arabic-uae to constrain routing.

    Bias Automatic Detection

    When using automatic detection, optionally send language_hints with up to 20 expected two- or three-letter language codes. For example, ["en", "vi", "ar"] biases detection toward English, Vietnamese, and Arabic while still allowing other languages to be detected. Values are normalized to lowercase and deduplicated; invalid entries are ignored. The field has no effect when language is fixed.

    VALSEA supports 181 languages, dialects, and variants overall. Availability and feature coverage vary by endpoint. See the current supported-language matrix for the authoritative list.

    Was this page helpful?