# ds1000 / 923 - taskset: [ds1000](https://harnessreport.com/tasks/ds1000.md) - difficulty: - category: - language: - runnable from the site: no - agent timeout: 1800s ## Results by harness _none yet_ ## Instruction ``` # 923: DS-1000 Task ## Prompt Problem: I have a data which include dates in sorted order. I would like to split the given data to train and test set. However, I must to split the data in a way that the test have to be newer than the train set. Please look at the given example: Let's assume that we have data by dates: 1, 2, 3, ..., n. The numbers from 1 to n represents the days. I would like to split it to 20% from the data to be train set and 80% of the data to be test set. Good results: 1) train set = 1, 2, 3, ..., 20 test set = 21, ..., 100 2) train set = 101, 102, ... 120 test set = 121, ... 200 My code: train_size = 0.2 train_dataframe, test_dataframe = cross_validation.train_test_split(features_dataframe, train_size=train_size) train_dataframe = train_dataframe.sort(["date"]) test_dataframe = test_dataframe.sort(["date"]) Does not work for me! Any suggestions? A: <code> import numpy as np import pandas as pd from sklearn.model_selection import train_test_split features_dataframe = load_data() def solve(features_dataframe): # return the solution in this function # train_dataframe, test_dataframe = solve(features_dataframe) ### BEGIN SOLUTION ## What to do - Edit `solution/solution.py` so the code passes the DS-1000 tests. - Do not access the internet or install new packages; required libraries are preinstalled in the Docker image. - Run tests locally via `bash tests/test.sh`. ## Notes - Keep the variable names/signatures implied by the prompt/code_context. - The evaluator uses the original DS-1000 `code_context` (`test_execution` / `test_string`). ``` --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp