# ds1000 / 922 - taskset: [ds1000](https://harnessreport.com/tasks/ds1000.md) - difficulty: - category: - language: - runnable from the site: no - agent timeout: 1800s ## Results by harness _none yet_ ## Instruction ``` # 922: DS-1000 Task ## Prompt Problem: I have a data which include dates in sorted order. I would like to split the given data to train and test set. However, I must to split the data in a way that the test have to be older than the train set. Please look at the given example: Let's assume that we have data by dates: 1, 2, 3, ..., n. The numbers from 1 to n represents the days. I would like to split it to 80% from the data to be train set and 20% of the data to be test set. Good results: 1) train set = 21, ..., 100 test set = 1, 2, 3, ..., 20 2) train set = 121, ... 200 test set = 101, 102, ... 120 My code: train_size = 0.8 train_dataframe, test_dataframe = cross_validation.train_test_split(features_dataframe, train_size=train_size) train_dataframe = train_dataframe.sort(["date"]) test_dataframe = test_dataframe.sort(["date"]) Does not work for me! Any suggestions? A: <code> import numpy as np import pandas as pd from sklearn.model_selection import train_test_split features_dataframe = load_data() </code> train_dataframe, test_dataframe = ... # put solution in these variables BEGIN SOLUTION <code> ## What to do - Edit `solution/solution.py` so the code passes the DS-1000 tests. - Do not access the internet or install new packages; required libraries are preinstalled in the Docker image. - Run tests locally via `bash tests/test.sh`. ## Notes - Keep the variable names/signatures implied by the prompt/code_context. - The evaluator uses the original DS-1000 `code_context` (`test_execution` / `test_string`). ``` --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp