# Speech-to-Text API

### Powered by Whisper v3 - Convert audio to text quickly and reliably.

### Speaker diarization - Automatically detect who is speaking.

Just $0.50 per 3 hours of speech - Lowest price on the market.

Use our audio-to-text API to build AI-powered features such as automatically generated subtitles, summaries of podcasts, or audio chats. Our API uses the latest Whisper large-v3 AI model to deliver accurate transcriptions with minimal latency and the most competitive pricing available. Transcribe 30 minutes of audio in under one minute. More than 100 languages are supported.

### API Usage

Our OpenAI compatible API makes it easy to switch. If you haven't already, you will need to [create an API key](/content/apis/keys/index.html) to authenticate your requests.

Use OpenAI library

JavascriptPythonCurl

```
const body = new FormData();
body.append('file', 'https://output.lemonfox.ai/wikipedia_ai.mp3');
// instead of providing a URL you can also upload a file object:
// body.append('file', new Blob([await fs.readFile('/path/to/audio.mp3')]));
body.append('language', 'english');
body.append('response_format', 'json');

fetch('https://api.lemonfox.ai/v1/audio/transcriptions', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer YOUR_API_KEY'
  },
  body: body
})
.then(response => response.json()).then(data => {
  console.log(data['text']);
})
.catch(error => {
  console.error('Error:', error);
});
```

### API Response

Choose between different response formats to get the transcript in the format that best suits your needs. `VTT` and `SRT` are file formats that include timestamps and can be used to display subtitles in video players.

jsontextsrtverbose_jsonvtt

```
{"text": "Artificial intelligence is the intelligence of machines or software, as opposed to the intelligence of humans or animals. It is also the field of study in computer science that develops and studies intelligent machines."}
```

### API Parameters

The API `POST https://api.lemonfox.ai/v1/audio/transcriptions` takes the following parameters:

- **file**: File object / URL, required  
  The audio file to transcribe. You can either upload a file object to the API or provide a public URL to download the audio file. The upload size is limited to 100MB. When providing the audio file via URL the maximum file size is 1GB.  
  Supported audio and video file formats: `mp3`, `wav`, `flac`, `aac`, `opus`, `ogg`, `m4a`, `mp4`, `mpeg`, `mov`, `webm`, and more.

- **response_format**: string, optional, default: json  
  The format in which the generated transcript is returned. See example API responses above. Must be one of the following: `json`, `text`, `srt`, `verbose_json`, or `vtt`.

`vtt` and `srt` are file formats that can be used to add subtitles to videos. `verbose_json` contains additional information such as the duration of the audio and timestamps for each audio segment.

- **speaker_labels**: boolean, optional  
  Set this parameter to `true` to enable speaker diarization. This will add speaker labels to the transcript. Currently the maximum number of speakers is limited to 4.

**Important:** Make sure to set the `response_format` parameter to `verbose_json` to be able to access the speaker labels. You can find an example response above.

- **prompt**: string, optional  
  A text to guide the transcript's style or continue a previous audio transcript. The prompt should be in the same language as the audio.

- **language**: string, optional  
  The language of the input audio. If no language is provided we detect the language automatically. Supplying the input language can improve accuracy and latency.

- **callback_url**: URL, optional  
  A URL to which the API will send a POST request when the transcription is ready. The POST request will include the transcript in the specified response format.

- **translate**: boolean, optional  
  Set this parameter to `true` to translate the audio content to English.

- **timestamp_granularities[]**: array, optional  
  Enable word-level timestamps by adding `word` to the array (eg `timestamp_granularities[]=word`). By default only timestamps for each segment are added to the response. To use this feature, `response_format` must be set to `verbose_json`.

🇪🇺 EU-based processing

By default, API requests are processed by servers around the world. To enable EU-based processing use `eu-api.lemonfox.ai` instead of `api.lemonfox.ai` as the API endpoint. This ensures that your data is processed within the EU.

Note: EU-based processing incurs a 20% surcharge, i.e. the price for processing 3 hours of audio is $0.60 (instead of $0.50).
