2. Running available benchmarks
We have several benchmarks to run:
plasticc- simple ETL and ML for plasticc dataset https://plasticc.org/data-release/ny_taxi- 4 queries (mainly gropuby) for NY taxi dataset https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.pageny_taxi_ml- simple ETL and ML based on NY taxi dataset https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page
Each benchmark’s source code is stored in it’s own folder in timedf_benchmarks/
2.1. Running benchmark
Let’s run one of benchmarks (plasticc) starting from a system with installed timedf in conda environment named ENV_NAME="timedf".
Activate your conda environment:
export ENV_NAME="timedf" && conda activate $ENV_NAME.Download data using (instructions)[https://github.com/intel-ai/timedf/blob/master/DATASETS.md].
- Run benchmark with pandas:
benchmark-run plasticc -data_file ./datasets/plasticc -backend Pandas ${DB_COMMON_OPTS}. To run with with modin on ray replace
"Pandas"->"Modin_on_ray", for modin on HDK replace"Pandas"->"Modin_on_hdk".You can get a list of all possible parameters with
benchmark-run -h.Optinal, not needed for plasticc. You might need to install benchmark-specific dependencies with:
conda env -n $ENV_NAME update -f timedf_benchmarks/$BENCHMARK_NAME/requirements.yamlOptinal. If you want to store results in a database, define environment variable with parameters:
export DB_COMMON_OPTS="". For example, to save results to local sqlite database (essentially just file on your filesystem) useexport DB_COMMON_OPTS="-db_name db.sqlite"
- Run benchmark with pandas:
If you want to customize how this run is stored in the database use these arguments:
-save_benchmark_name BENCHMARK_NAME- benchmark name for DB storage-save_backend_name BACKEND_NAME- name of the backend used for DB storage. You can use this for experimental branches of libraries.-tag TAG- tag for this run. This will be stored intagcolumn. Useful for identifying unique runs, such as experimental library versions.
2.2. Validating intermediate dataframes
You might want to validate that dataframe processing library is providing results, consistent with other libraries. This is an optional result validation feature that is not yet available, but will be provided in the future.
TBD