Speech-to-Text API
Powered by Whisper v3 - Convert audio to text quickly and reliably.
Speaker diarization - Automatically detect who is speaking.
Just $0.50 per 3 hours of speech - Lowest price on the market.
Use our audio-to-text API to build AI-powered features such as automatically generated subtitles, summaries of podcasts, or audio chats. Our API uses the latest Whisper large-v3 AI model to deliver accurate transcriptions with minimal latency and the most competitive pricing available. Transcribe 30 minutes of audio in under one minute. More than 100 languages are supported.
API Usage
Our OpenAI compatible API makes it easy to switch. If you haven't already, you will need to create an API key to authenticate your requests.
Use OpenAI library
JavascriptPythonCurl
const body = new FormData();
body.append('file', 'https://output.lemonfox.ai/wikipedia_ai.mp3');
// instead of providing a URL you can also upload a file object:
// body.append('file', new Blob([await fs.readFile('/path/to/audio.mp3')]));
body.append('language', 'english');
body.append('response_format', 'json');
fetch('https://api.lemonfox.ai/v1/audio/transcriptions', {
method: 'POST',
headers: {
'Authorization': 'Bearer YOUR_API_KEY'
},
body: body
})
.then(response => response.json()).then(data => {
console.log(data['text']);
})
.catch(error => {
console.error('Error:', error);
});
API Response
Choose between different response formats to get the transcript in the format that best suits your needs. VTT and SRT are file formats that include timestamps and can be used to display subtitles in video players.
jsontextsrtverbose_jsonvtt
{"text": "Artificial intelligence is the intelligence of machines or software, as opposed to the intelligence of humans or animals. It is also the field of study in computer science that develops and studies intelligent machines."}
API Parameters
The API POST https://api.lemonfox.ai/v1/audio/transcriptions takes the following parameters:
file: File object / URL, required
The audio file to transcribe. You can either upload a file object to the API or provide a public URL to download the audio file. The upload size is limited to 100MB. When providing the audio file via URL the maximum file size is 1GB.
Supported audio and video file formats:mp3,wav,flac,aac,opus,ogg,m4a,mp4,mpeg,mov,webm, and more.response_format: string, optional, default: json
The format in which the generated transcript is returned. See example API responses above. Must be one of the following:json,text,srt,verbose_json, orvtt.
vtt and srt are file formats that can be used to add subtitles to videos. verbose_json contains additional information such as the duration of the audio and timestamps for each audio segment.
- speaker_labels: boolean, optional
Set this parameter totrueto enable speaker diarization. This will add speaker labels to the transcript. Currently the maximum number of speakers is limited to 4.
Important: Make sure to set the response_format parameter to verbose_json to be able to access the speaker labels. You can find an example response above.
prompt: string, optional
A text to guide the transcript's style or continue a previous audio transcript. The prompt should be in the same language as the audio.language: string, optional
The language of the input audio. If no language is provided we detect the language automatically. Supplying the input language can improve accuracy and latency.callback_url: URL, optional
A URL to which the API will send a POST request when the transcription is ready. The POST request will include the transcript in the specified response format.translate: boolean, optional
Set this parameter totrueto translate the audio content to English.timestamp_granularities[]: array, optional
Enable word-level timestamps by addingwordto the array (egtimestamp_granularities[]=word). By default only timestamps for each segment are added to the response. To use this feature,response_formatmust be set toverbose_json.
🇪🇺 EU-based processing
By default, API requests are processed by servers around the world. To enable EU-based processing use eu-api.lemonfox.ai instead of api.lemonfox.ai as the API endpoint. This ensures that your data is processed within the EU.
Note: EU-based processing incurs a 20% surcharge, i.e. the price for processing 3 hours of audio is $0.60 (instead of $0.50).