Qwen vision-language model for visual reasoning, documents, and agent tasks
Details
| Field | Value |
|---|---|
| id | Qwen/Qwen2.5-VL-72B-Instruct |
| modality | text+image->text |
| context_length | 128000 |
| max_input_tokens | 120000 |
| max_output_tokens | 8192 |
| quantization | fp8 |
| knowledge_cutoff | 2024-12 |
| release_date | 2025-01-20 |
| is_open_weights | yes |
| supports_tool_call | yes |
| supports_reasoning | no |
| supports_structured_output | yes |
| supports_attachment | yes |
| input_modalities | text, image |
| output_modalities | text |