Benchmark: Micro benchmark - Add float datatype support and other refinements to GPU Stream #769

WenqingLan1 · 2025-12-19T20:05:13Z

Refinements:

Use 128-bit aligned vector types (double2/float4) for optimal memory bandwidth.
Add support for float execution.
Add --data_type <float|double> CLI option for runtime type selection.
Move template kernel implementations to header file (required for CUDA template instantiation across compilation units).
Rename entry point file from gpu_stream_test.cpp to gpu_stream_main.cpp.
Updated hard-coded GPU iteration to single node run so it can run with SuperBench's distributed execution in config.yaml.
Updated numa assignment from hard coded numa_alloc_onnode to numa_alloc_local to optimize memory allocation.
Updated micro benchmark doc to reflect new metric name removing gpu_id.

New config:

    gpu-stream:fp64:
      <<: *default_local_mode
      timeout: 600
      parameters:
        num_warm_up: 10
        num_loops: 40
        size: 1308622848
        data_type: double
    gpu-stream:fp64-correctness:
      <<: *default_local_mode
      timeout: 600
      parameters:
        num_warm_up: 0
        num_loops: 1
        size: 1048576
        data_type: double
        check_data: true
    gpu-stream:fp32:
      <<: *default_local_mode
      timeout: 600
      parameters:
        num_warm_up: 10
        num_loops: 40
        size: 2617245696
        data_type: float
    gpu-stream:fp32-correctness:
      <<: *default_local_mode
      timeout: 600
      parameters:
        num_warm_up: 0
        num_loops: 1
        size: 1048576
        data_type: float
        check_data: true

New rule:

    gpu-stream:
      statistics:
        - mean
      categories: GPU-STREAM
      aggregate: True
      metrics:
        - gpu-stream:fp(?:32|64)/STREAM_.*_(?:bw|ratio):(\d+)

Example results:

"gpu-stream:fp32/STREAM_COPY_float_buffer_2617245696_block_256_bw:0": 1234, 
"gpu-stream:fp32/STREAM_COPY_float_buffer_2617245696_block_256_bw:1": 1234, 
"gpu-stream:fp32/STREAM_COPY_float_buffer_2617245696_block_256_bw:2": 1234, 
"gpu-stream:fp32/STREAM_COPY_float_buffer_2617245696_block_256_bw:3": 1234

Processed by rules:

| gpu-stream:fp32/STREAM_COPY_float_buffer_2617245696_block_256_bw | mean | 1234|

codecov · 2025-12-19T20:14:11Z

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 85.70%. Comparing base (575859b) to head (d8a91ab).

Additional details and impacted files

@@           Coverage Diff           @@
##             main     #769   +/-   ##
=======================================
  Coverage   85.70%   85.70%           
=======================================
  Files         102      102           
  Lines        7703     7704    +1     
=======================================
+ Hits         6602     6603    +1     
  Misses       1101     1101

Flag	Coverage Δ
cpu-python3.10-unit-test	`70.96% <50.00%> (+<0.01%)`	⬆️
cpu-python3.12-unit-test	`70.96% <50.00%> (+<0.01%)`	⬆️
cpu-python3.7-unit-test	`70.44% <50.00%> (+<0.01%)`	⬆️
cuda-unit-test	`83.59% <100.00%> (+<0.01%)`	⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:

❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Copilot

Pull request overview

Updates the GPU STREAM microbenchmark to support runtime-selectable FP32/FP64 execution and improve GPU memory bandwidth utilization, while aligning SuperBench integration (CLI, output tags, docs, and tests) to the new behavior.

Changes:

Add --data_type <float|double> to select FP32/FP64 at runtime and propagate it through the Python benchmark wrapper + unit tests.
Refactor CUDA kernels to use 128-bit vectorized accesses (double2 / float4) and move template kernel implementations into a header for cross-TU instantiation.
Adjust execution/output to single visible GPU (device 0 via CUDA_VISIBLE_DEVICES) and update metric/tag formats (removing gpu_id) plus docs/examples/test log.

Reviewed changes

Copilot reviewed 11 out of 13 changed files in this pull request and generated 5 comments.

Show a summary per file

File	Description
`tests/data/gpu_stream.log`	Updates golden log output to include data type and new tag format (no `gpu_id`).
`tests/benchmarks/micro_benchmarks/test_gpu_stream.py`	Extends command-generation assertions to include `--data_type` (currently only covers `double`).
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream_utils.hpp`	Removes NUMA/GPU iteration fields from args and adds `Opts::data_type`.
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream_utils.cpp`	Adds CLI parsing/printing for `--data_type`.
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream_main.cpp`	New entry point replacing the previous main file.
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream_kernels.hpp`	Introduces vector-type mapping and templated kernel definitions (128-bit loads/stores).
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream_kernels.cu`	Keeps a CUDA compilation unit and moves template implementations to the header.
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.hpp`	Expands bench-args variant to support `float` and `double`.
`superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu`	Uses local NUMA allocation, enforces 16B/thread sizing, launches templated vectorized kernels, updates tag format, and runs only CUDA device 0.
`superbench/benchmarks/micro_benchmarks/gpu_stream/CMakeLists.txt`	Switches target sources to the new `gpu_stream_main.cpp`.
`superbench/benchmarks/micro_benchmarks/gpu_stream.py`	Adds `--data_type` argument and forwards it to the binary.
`examples/benchmarks/gpu_stream.py`	Updates example invocation to include `--data_type double`.
`docs/user-tutorial/benchmarks/micro-benchmarks.md`	Updates gpu-stream metric patterns to include `(double

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu

docs/user-tutorial/benchmarks/micro-benchmarks.md

Copilot

Pull request overview

Copilot reviewed 12 out of 14 changed files in this pull request and generated 2 comments.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu

WenqingLan1 added 3 commits December 18, 2025 01:16

remove fixed gpu id & numa id assignment

7f23c75

use 128bit alignment, add float support, cleanup

d63fe8c

add data_type arg

242714e

WenqingLan1 requested a review from a team as a code owner December 19, 2025 20:05

WenqingLan1 added the micro-benchmarks Micro Benchmark Test for SuperBench Benchmarks label Dec 19, 2025

guoshzhao self-assigned this Dec 19, 2025

guoshzhao requested review from guoshzhao and polarG December 19, 2025 20:32

WenqingLan1 and others added 5 commits December 19, 2025 23:31

fix lint

e8d0282

fix clang lint

5a18946

update doc

fddf56e

Merge branch 'main' into wenqinglan/refine-gpu-stream

3c359a3

Merge branch 'microsoft:main' into wenqinglan/refine-gpu-stream

e445363

Copilot AI review requested due to automatic review settings February 3, 2026 22:14

Copilot started reviewing on behalf of WenqingLan1 February 3, 2026 22:15 View session

Copilot AI reviewed Feb 3, 2026

View reviewed changes

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu Outdated Show resolved Hide resolved

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu Outdated Show resolved Hide resolved

docs/user-tutorial/benchmarks/micro-benchmarks.md Outdated Show resolved Hide resolved

WenqingLan1 and others added 2 commits February 5, 2026 16:04

Merge branch 'microsoft:main' into wenqinglan/refine-gpu-stream

60b130c

fix alloc count & comment

f31933f

Copilot AI review requested due to automatic review settings February 6, 2026 00:20

Copilot AI reviewed Feb 6, 2026

View reviewed changes

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu Show resolved Hide resolved

superbench/benchmarks/micro_benchmarks/gpu_stream/gpu_stream.cu Show resolved Hide resolved

fix: reset gpu-burn submodule to correct commit

d8a91ab

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Benchmark: Micro benchmark - Add float datatype support and other refinements to GPU Stream #769

Benchmark: Micro benchmark - Add float datatype support and other refinements to GPU Stream #769

Uh oh!

WenqingLan1 commented Dec 19, 2025 •

edited

Loading

Uh oh!

codecov bot commented Dec 19, 2025 •

edited

Loading

Uh oh!

Copilot AI left a comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Copilot AI left a comment

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Benchmark: Micro benchmark - Add float datatype support and other refinements to GPU Stream #769

Are you sure you want to change the base?

Benchmark: Micro benchmark - Add float datatype support and other refinements to GPU Stream #769

Uh oh!

Conversation

WenqingLan1 commented Dec 19, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

codecov bot commented Dec 19, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Codecov Report

Uh oh!

Copilot AI left a comment

Choose a reason for hiding this comment

Pull request overview

Reviewed changes

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Copilot AI left a comment

Choose a reason for hiding this comment

Pull request overview

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

WenqingLan1 commented Dec 19, 2025 •

edited

Loading

codecov bot commented Dec 19, 2025 •

edited

Loading