Scripts to run presto in a Slurm cluster by misiugodfrey · Pull Request #126 · rapidsai/velox-testing

misiugodfrey · 2025-11-13T02:24:04Z

Creating a set of scripts to automate running presto (gpu-native) within a Slurm cluster. Currently should support single-node-single-gpu and single-node-multi-gpu (multiple workers on the same node with different GPU assignments). It's currently being used to run TPCH-SF1k

TODOs:
Refactor this to use more of the velox-testing scripts internally (such as run_benchmarks.sh).
Create a common config setup

misiugodfrey · 2025-11-17T18:15:12Z

@@ -0,0 +1,45 @@
+#!/bin/bash
+#SBATCH --time=00:25:00


This can take a little while due to the cost of the ANALYZE steps. Currently 25 min seems to be a decent limit.

misiugodfrey · 2025-11-17T18:44:13Z

+sbatch "$@" --nodes=1 --ntasks-per-node=10 ${JOB_TYPE}_benchmarks.sbatch;
+
+echo "Waiting for jobs to finish..."
+while :; do


This just tracks the job in the terminal until it is finished.

paul-aiyedun

The overall approach makes sense to me. However, I had a few questions and code cleanup comments.

paul-aiyedun · 2025-11-20T18:53:12Z

Can we give this file a more specific name e.g. echo_helper.sh?

paul-aiyedun · 2025-11-20T18:58:49Z

+rm *.log
+rm *.out
+
+[ $# -ge 1 ] && echo "$0 expected first argument is 'create/run'" && exit 1


Nit: I think using if statements for this type of check is more readable.

paul-aiyedun · 2025-11-20T19:01:49Z

+
+[ $# -ge 1 ] && echo "$0 expected first argument is 'create/run'" && exit 1
+JOB_TYPE="$1"
+[ "$JOB_TYPE" == "create" ] && [ "$JOB_TYPE" == "run" ] && echo "parameter must be create or run" && exit 1


Can we create different scripts for each workflow?

paul-aiyedun · 2025-11-20T19:02:52Z

+[ "$JOB_TYPE" == "create" ] && [ "$JOB_TYPE" == "run" ] && echo "parameter must be create or run" && exit 1
+shift 1
+
+[ -z "$NUM_NODES" ] && echo "NUM_NODES env variable must be set" && exit 1


Can we consistently pass these variables as script arguments instead of a mix of environment variables and arguments?

paul-aiyedun · 2025-11-20T19:25:54Z

+        fi
+    done
+    if ((${#missing[@]})); then
+        echo_error "required env var ${missing[*]} not set"


Should this file source common.sh?

paul-aiyedun · 2025-11-20T20:26:36Z

+source slurm_functions.sh
+source setup_coord.sh
+
+for i in $(seq 0 $(( $NUM_WORKERS - 1 )) ); do


Is this script run per node or per task?

paul-aiyedun · 2025-11-20T21:11:32Z

+    rm ${WORKSPACE}/iterating_queries.sql
+}
+
+# Check if the coordinator is running via curl.  Fail after 10 retries.


Can we reuse the existing wait_for_worker_node_registration function (

velox-testing/presto/scripts/common_functions.sh

Line 17 in b49c408

function wait_for_worker_node_registration() {

)?

paul-aiyedun · 2025-11-20T21:13:01Z

+    validate_environment_preconditions CONFIGS SINGLE_NODE_EXECUTION
+    local coord_config="${CONFIGS}/etc_coordinator/config_native.properties"
+    # Replace placeholder in configs
+    sed -i "s+discovery\.uri.*+discovery\.uri=http://${COORD}:8080+g" ${coord_config}


I think we should expand our config generation capability in velox-testing and avoid having to use sed commands.

paul-aiyedun · 2025-11-20T21:19:21Z

+    # so that the job will finish when the cli is done (terminating background
+    # processes like the coordinator and workers).
+    if [ "${type}" == "coord" ]; then
+        srun -w $COORD --ntasks=1 --overlap \


Is this kicking off another job or running the container locally?

paul-aiyedun · 2025-11-20T21:34:38Z

I think it might make sense to name the top level directory slurm instead of cluster.

mattgara · 2025-11-21T00:39:50Z

+    mkdir -p ${WORKSPACE}/.hive_metastore
+
+    # Run the worker with the new configs.
+    CUDA_VISIBLE_DIVICES=${gpu_id} srun -N1 -w $node --ntasks=1 --overlap \


Typo I believe

Untested cluster scripts to run presto in a Slurm cluster

77e7713

misiugodfrey requested a review from simoneves November 13, 2025 02:27

misiugodfrey and others added 9 commits November 12, 2025 18:32

localhost

38aceed

removed queued time

7462add

Tested version that re-uses configs

f6cb425

update velox-testing path

4a4a889

Merge branch 'main' into misiug/ClusterScripts

3fac0a6

update dispatch.sh

56c7582

Add readme and some outlining

b56e6b4

removed unnecessary file

1cf0018

removed dead code

0168e0b

misiugodfrey commented Nov 17, 2025

View reviewed changes

misiugodfrey marked this pull request as ready for review November 17, 2025 18:49

misiugodfrey requested a review from paul-aiyedun November 17, 2025 19:06

misiugodfrey changed the title ~~Untested cluster scripts to run presto in a Slurm cluster~~ Scripts to run presto in a Slurm cluster Nov 17, 2025

misiugodfrey added 4 commits November 18, 2025 16:02

Added multi-node support

690bf7d

fixes

d7d43af

overlap

dab8a03

update readme

a0cc453

paul-aiyedun reviewed Nov 20, 2025

View reviewed changes

mattgara reviewed Nov 21, 2025

View reviewed changes

misiugodfrey added 2 commits November 20, 2025 23:24

Added data generation to create option

ad33928

update README

75fef44

misiugodfrey marked this pull request as draft November 21, 2025 07:26

fix typo

a1ccc6e

Conversation

misiugodfrey commented Nov 13, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

paul-aiyedun left a comment

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

misiugodfrey commented Nov 13, 2025 •

edited

Loading