Multimodal model for analyzing text, images, documents, and rich media
Details
| Field | Value |
|---|---|
| id | bytedance/ui-tars-1.5-7b |
| modality | image+text->text |
| context_length | 128000 |
| max_output_tokens | 2048 |
| release_date | 2025-07-22 |
| is_open_weights | no |
| supports_tool_call | no |
| supports_reasoning | no |
| supports_structured_output | yes |
| supports_attachment | yes |
| input_modalities | image, text |
| output_modalities | text |