{"task": {"agent_timeout": 1800, "task": "widesearch-ws-en-090", "verifier_timeout": 600, "instruction": "I want to check the recent performance of large models. Please list Gemini 2.5 Pro & Lite & Flash-lite (separate the thinking and non-thinking models), Claude 3.7 and subsequently released models until June, 2025, O3 & O3 Mini & O4 Mini, Doubao-1.5-Thinking and 1.6-Thinking, DeepSeek V3 and subsequent main released models until June 2025. The specific metrics I want to know include the model number (such as Gemini-2.5-pro), context window (such as 32K), AIME-2025 metrics, SWE verified metrics (single attempt), and tau-bench-retail and airline metrics. Please try to search for all metrics on the official website of the corresponding model, but do not fabricate them. If the official website/official paper does not release them, the relevant metrics can be output as \"NA\".\n\nPlease output the sorted data in the format of one Markdown table. The column names in the table are: model name, company, context window, AIME 2025, SWE-bench Verified, TAU-bench-Airline, Tau-bench-Retail.\nDon't ask me any questions, just output the results according to the columns without omitting cells arbitrarily. The output format is ```markdown\n{data_content}\n```\n\n---\n\n**Important:** Write your final answer as a Markdown table to the file `/workspace/output.md`. The file should contain ONLY the Markdown table (with header row and separator row). Do not include any other text, explanation, or code blocks \u2014 just the raw Markdown table.\n", "memory": "4096m", "runnable": false, "difficulty": "hard", "language": "", "cpus": 1, "instruction_truncated": false, "category": "information-retrieval", "compose": false, "has_solution": true, "oracle": null, "docker_image": "", "taskset": "widesearch", "tags": []}, "runs": []}