Speech transcription model for accurate audio-to-text and captioning workflows
Details
| Field | Value |
|---|---|
| id | gpt-4o-transcribe |
| family | gpt |
| modality | text+audio->text |
| context_length | 16000 |
| max_output_tokens | 16000 |
| release_date | 2025-03-20 |
| is_open_weights | no |
| supports_tool_call | yes |
| supports_reasoning | no |
| supports_structured_output | no |
| supports_attachment | yes |
| input_modalities | text, audio |
| output_modalities | text |