feat: multi-gpu inference, trajectory analyzer #2
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Most of the PR is the trajectory analyzer, but this also includes the changes to the dependencies that made it possible to run inference outisde of apptainer and on the full 8 GPU node with no issues from vllm coming up.
@BjarniHaukur you should be able to just
uv syncand run it. See benchmarks/swe_bench/run_harness_eval.sh for the batchscript.I created a pyproject from scratch since I was facing some version conflicts, but haven't added training dependencies back to it. Let me know if adding them to the new setup is not enough, we can debug if that's not the case. After that we can make a run to see if the parallelism is also working for training.