Storage Benchmark Kit
PerlBench is SBK’s synthetic driver for measuring the end-to-end overhead
and scalability of the SBK harness and its PerL timestamp pipeline. It removes
storage, network, serialization, and vendor-SDK work from the operation path:
This is not a storage benchmark. Use it to compare PerL configurations, estimate the measurement ceiling of a machine, and check whether timestamp collection can keep up with a proposed storage workload.
Both drivers avoid external storage, but they test different parts of SBK:
| Property | PerlBench |
Null |
|---|---|---|
| Default write | Immediate synchronous completion | Incomplete future that remains pending until timeout or shutdown |
| Read | Immediate, preallocated payload | Sleeps for the configured timeout and returns null |
| Timestamp production | One completed measurement per operation | No completed measurement on the default path |
| Main question | How fast can SBK and PerL process completed operations? | Does idle reporting, timeout, interruption, and shutdown work correctly? |
Using Null for queue throughput would mostly measure an idle or blocked
worker. Using PerlBench for timeout behavior would never create the required
pending operation. Keeping separate driver identities prevents these results
from being confused.
flowchart LR
W["SBK worker"]
S["start = clock"]
OP["PerlBench no-op write<br/>or preallocated read"]
E["end = clock"]
Q["Timestamp queue"]
R["Single PerL recorder"]
O["Window and total results"]
W --> S --> OP --> E --> Q --> R --> O
classDef timed fill:#dcfce7,stroke:#166534,color:#111
classDef transport fill:#dbeafe,stroke:#1d4ed8,color:#111
class OP timed
class Q,R transport
The reported operation latency is end - start. Queue insertion occurs after
the end timestamp, so displayed latency percentiles do not directly measure
enqueue latency. Queue implementation affects:
Use the PerL JMH task when the research question is specifically queue add/poll latency and normalized bytes allocated per operation:
./gradlew :perl:timeStampQueuePerformanceTest
Use the PerlBench functional tasks when the question is whether that queue advantage remains visible through the complete SBK worker, clock, queue, recorder, and logger pipeline:
# Validate and measure each complete SBK path independently
./gradlew :drivers:perlbench:perlBenchMpscPerformanceTest
./gradlew :drivers:perlbench:perlBenchJdkClqPerformanceTest
# Run both, classify throughput, and verify allocation and correctness
./gradlew :drivers:perlbench:perlBenchQueuePerformanceTest
These are intentionally separate from ./gradlew check. The individual
tasks execute seven independently warmed JVM forks. The comparison randomizes
the queue order and uses a test-only maxQs=1 resource, so all writers publish
to one shared queue and the functional workload is genuinely MPSC. Each mode
must complete every requested record with zero invalid latencies.
The comparison reports the median, a 95% confidence interval for the mean, and one of three throughput classifications:
MPSC_FASTER when the MPSC interval is entirely above the JDK interval;JDK_CLQ_FASTER when the JDK interval is entirely above the MPSC interval;INCONCLUSIVE when the intervals overlap.Throughput classification is deliberately not a build gate because CPU scheduling, frequency scaling, virtualization, and concurrent host activity can change it. Exact counts, zero invalid latencies, and the structural allocation reduction remain hard gates. The task runs the queue JMH benchmark and includes its measured normalized allocation in the final comparison.
The workload can be adjusted without modifying source:
./gradlew :drivers:perlbench:perlBenchQueuePerformanceTest \
-PperlBenchPerformanceWriters=8 \
-PperlBenchPerformanceRecords=10000000 \
-PperlBenchPerformanceWarmupRecords=10000000 \
-PperlBenchPerformanceRuns=9
Reports are written to
drivers/perlbench/build/reports/perlbench-performance/. The displayed
operation latency is diagnostic only: SBK captures the operation end time
before enqueueing the timestamp. Therefore the functional comparison
characterizes end-to-end completed-record throughput, while the JMH result
provides the direct queue-latency and normalized-allocation measurements.
The common -mpscqueue SBK option works with real storage drivers and is
especially useful with PerlBench.
| Option | Values and default | Meaning |
|---|---|---|
-mpscqueue |
true or false; property default |
true selects intrusive TimeStampMpscQueue; false selects JDK ConcurrentLinkedQueue |
Queue topology is deliberately not exposed through the CLI. SBK loads
maxQs and qPerWorker from
sbk-api/src/main/resources/sbk.properties. The default maxQs=0 creates
per-worker channels and represents SBK’s normal high-throughput deployment.
Advanced queue research can set maxQs=1 and rebuild SBK to force every
writer or reader to publish to one shared MPSC queue.
At startup SBK prints the effective selection:
PerL Timestamp Queue: TimeStampMpscQueue (MPSC)
PerL Timestamp Queue Topology: 10 queue(s) per worker
or:
PerL Timestamp Queue: ConcurrentLinkedQueue (JDK)
PerL Timestamp Queue Topology: 10 queue(s) per worker
-records without -seconds divides the requested total across the workers
and waits until every assigned operation has completed and every timestamp has
been consumed. Odd totals are supported; the first workers receive the
remainder one record at a time.
Compare the two queue paths with the same command except for -mpscqueue:
# Intrusive timestamp node and queue
./build/install/sbk/bin/sbk \
-class perlbench -writers 4 -size 1024 -records 20000000 \
-time ns -thread p -mpscqueue true
# TimeStamp plus the JDK queue's internal node
./build/install/sbk/bin/sbk \
-class perlbench -writers 4 -size 1024 -records 20000000 \
-time ns -thread p -mpscqueue false
Confirm that each result reports exactly 20000000 records and
0 invalid latencies. Record mode is the best end-to-end check for lost,
duplicated, or undrained measurements.
-seconds runs at maximum speed when -throughput remains -1. At the
duration boundary SBK stops workload generation and reports the measurements
accumulated by that hard deadline. Timed mode answers: “How many operations
can this complete pipeline sustain for this period?”
./build/install/sbk/bin/sbk \
-class perlbench -writers 1 -size 1024 -seconds 30 \
-time ns -thread p -mpscqueue true
./build/install/sbk/bin/sbk \
-class perlbench -writers 1 -size 1024 -seconds 30 \
-time ns -thread p -mpscqueue false
Repeat with -writers 2, 4, 8, and the available CPU count. A queue can
win with one producer yet lose at a particular contention level, so one thread
count is not a scalability study.
Use -throughput MBPS with -seconds to hold offered data throughput near a
target. This separates behavior at a realistic load from maximum-saturation
behavior.
For 1,024-byte records, -throughput 1000 is approximately 1,048,576
records/second across all workers:
./build/install/sbk/bin/sbk \
-class perlbench -writers 4 -size 1024 -seconds 30 -throughput 1000 \
-time ns -thread p -mpscqueue true
./build/install/sbk/bin/sbk \
-class perlbench -writers 4 -size 1024 -seconds 30 -throughput 1000 \
-time ns -thread p -mpscqueue false
Both runs should approach the same configured MB/s. Compare high-percentile latencies, CPU consumption, allocation, GC pauses, and queue-drain stability. Do not interpret a small difference in the displayed no-op latency as direct queue latency because enqueue is outside the timed interval.
The reader reuses one payload per reader worker, so it exercises SBK’s reader control flow without per-record driver allocation:
./build/install/sbk/bin/sbk \
-class perlbench -readers 4 -size 1024 -records 1000000 \
-time ns -thread p -mpscqueue true
Run the same command with -mpscqueue false. No preparation step is required.
./gradlew clean :pathingJar installDist --rerun-tasks.jcmd,
/usr/bin/time -v, or equivalent OS tools.maxQs=1 and the
default maxQs=0 in sbk.properties.The following results are observations, not performance guarantees. They were collected on 2026-07-27 on a 16-vCPU VMware guest with an Intel Xeon Platinum 8462Y+, Linux 5.15, JDK 25.0.2, ZGC, and compact object headers.
| Metric | TimeStampMpscQueue |
JDK queue path | Observed difference |
|---|---|---|---|
| Add/poll round trip | 32.732 ns/op | 47.000 ns/op | MPSC 30.36% lower |
| Normalized allocation | 40 B/op | 56 B/op | MPSC 28.57% lower |
| Four-producer enqueue throughput | 5.85 M ops/s | 5.48 M ops/s | MPSC 6.74% higher; 99.9% confidence intervals overlapped |
The allocation difference is structural: the intrusive
TimeStampNode is both payload and link, whereas the JDK path allocates a
TimeStamp plus a private queue node. Throughput remains dependent on
contention, scheduling, and consumer service rate. The canonical exact values,
confidence intervals, environment, and interpretation are maintained in
the queue research guide;
this summary must not be interpreted as a universal throughput ranking.
| Workload | TimeStampMpscQueue |
JDK queue path | Interpretation |
|---|---|---|---|
| One writer, six-second saturation | 7.87 M records/s | 6.20 M records/s | MPSC 26.9% higher in this single run |
| Four writers, six-second saturation | 5.61 M records/s | 4.47 M records/s | MPSC 25.5% higher in this single run |
| Four writers, 1,000 MB/s target | 993.60 MB/s | 993.35 MB/s | Both sustained the requested load |
| One reader, six-second saturation | 8.11 M records/s | 6.16 M records/s | MPSC 31.6% higher in this single run |
| Four writers, 20 M exact records, three-run median | 5.53 M records/s | 5.19 M records/s | MPSC 6.6% higher under shared-queue contention |
| Same exact-record test, median peak RSS | 1.24 GiB | 1.54 GiB | MPSC used about 20% less process memory |
Every exact-count run completed all requested records with zero invalid latencies. The results demonstrate the expected memory advantage and show that adopting JDK CLQ’s one-node tail-slack strategy removed unnecessary shared-tail updates without sacrificing the intrusive queue’s latency. Results remain environment-specific, so the fallback and runtime comparison options remain valuable.
For queue algorithm diagrams, memory-ordering details, correctness evidence, and the complete JMH methodology, see TimeStampMpscQueue: architecture, correctness, and performance.