{"task": {"agent_timeout": 3000, "task": "pandas-dev__pandas-56523", "verifier_timeout": 6000, "instruction": "PERF: merge_ordered(how=\"inner\") without prior filtering is slow and uses much more memory.\n### Pandas version checks\n\n- [X] I have checked that this issue has not already been reported.\n\n- [X] I have confirmed this issue exists on the [latest version](https://pandas.pydata.org/docs/whatsnew/index.html) of pandas.\n\n- [ ] I have confirmed this issue exists on the main branch of pandas.\n\n\n### Reproducible Example\n\n```python\nimport pandas as pd\nimport numpy as np\n\nstart, stop, step = 0, 100_000_000, 2\ndf1 = pd.DataFrame({\"key\": np.arange(start, stop, step), \"val1\": np.arange(start, stop, step), \"val2\": np.arange(start, stop, step)}, copy=False)\ndf1.info()\n\nstart, stop, step = 50_000_000, 70_000_000, 1\ndf2 = pd.DataFrame({\"key\": np.arange(start, stop, step), \"val3\": np.arange(start, stop, step), \"val4\": np.arange(start, stop, step)}, copy=False)\ndf2.info()\n\n# This filtering should not be necessary for inner join as it only drops unused data\ndf1 = df1.query(\"key >= @df2.key.min() and key <= @df2.key.max()\")\ndf2 = df2.query(\"key in @df1.key\")\n\ndf = pd.merge_ordered(df1, df2, on=\"key\", how=\"inner\")\ndf.info()\n```\n\n### Installed Versions\n\n<details>\n\n```\nINSTALLED VERSIONS\n------------------\ncommit              : 2a953cf80b77e4348bf50ed724f8abc0d814d9dd\npython              : 3.10.9.final.0\npython-bits         : 64\nOS                  : Linux\nOS-release          : 6.2.0-36-generic\nVersion             : #37~22.04.1-Ubuntu SMP PREEMPT_DYNAMIC Mon Oct  9 15:34:04 UTC 2\nmachine             : x86_64\nprocessor           : x86_64\nbyteorder           : little\nLC_ALL              : None\nLANG                : pl_PL.UTF-8\nLOCALE              : pl_PL.UTF-8\n\npandas              : 2.1.3\nnumpy               : 1.26.2\npytz                : 2023.3.post1\ndateutil            : 2.8.2\nsetuptools          : 65.5.0\npip                 : 22.3.1\nCython              : None\npytest              : None\nhypothesis          : None\nsphinx              : None\nblosc               : None\nfeather             : None\nxlsxwriter          : None\nlxml.etree          : None\nhtml5lib            : None\npymysql             : None\npsycopg2            : None\njinja2              : None\nIPython             : None\npandas_datareader   : None\nbs4                 : None\nbottleneck          : None\ndataframe-api-compat: None\nfastparquet         : None\nfsspec              : None\ngcsfs               : None\nmatplotlib          : 3.8.2\nnumba               : None\nnumexpr             : None\nodfpy               : None\nopenpyxl            : None\npandas_gbq          : None\npyarrow             : None\npyreadstat          : None\npyxlsb              : None\ns3fs                : None\nscipy               : None\nsqlalchemy          : None\ntables              : None\ntabulate            : None\nxarray              : None\nxlrd                : None\nzstandard           : None\ntzdata              : 2023.3\nqtpy                : None\npyqt5               : None\n```\n\n</details>\n\n\n### Prior Performance\n\n![diff](https://github.com/pandas-dev/pandas/assets/8440636/530b6169-3892-44ea-bc79-5a25e067ccc0)\nPlot generated by `memory_profiler`:\n* black line: with 2 queries before `merge_ordered`\n* blue line: without 2 queries before `merge_ordered`\n", "memory": "8192m", "runnable": false, "difficulty": "hard", "language": "", "cpus": 1, "instruction_truncated": false, "category": "debugging", "compose": false, "has_solution": true, "oracle": null, "docker_image": "", "taskset": "swegym", "tags": ["debugging", "swe-bench"]}, "runs": []}