Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior
Details
| Field | Value |
|---|---|
| id | gpt-realtime-2.1 |
| family | gpt |
| modality | text+audio+image->text+audio |
| context_length | 128000 |
| max_input_tokens | 96000 |
| max_output_tokens | 32000 |
| knowledge_cutoff | 2024-09-30 |
| release_date | 2026-07-06 |
| is_open_weights | no |
| supports_tool_call | yes |
| supports_reasoning | yes |
| supports_structured_output | no |
| supports_attachment | yes |
| input_modalities | text, audio, image |
| output_modalities | text, audio |