Storage Benchmark Kit
This document explains how SBK’s MinIO driver turns a benchmark command into
S3 requests. It is intended for reviewers, maintainers, and performance
engineers who need to know exactly what a reported operation measures.
For operating procedures, use the ECS/OBS benchmark runbook. For the complete command-line reference and backend tutorials, use the driver README. Software agents must also follow the ECS/OBS agent workflow.
SBK writer/reader worker
|
v
MinIOWriter / MinIOReader
|
+-- S3OperationMix deterministic operation selection
+-- S3ObjectCatalog existing/published object selection
+-- S3RetryPolicy optional bounded retry sequence
+-- S3AsyncExecutor optional bounded SDK-future tracking
|
v
MinIO Java SDK 8.5.17
|
v
OkHttp connection pool / dispatcher
|
v
S3 endpoint: MinIO, ECS/ObjectScale, AWS S3, Ceph RGW, ...
The driver deliberately uses the vendor SDK instead of implementing S3 signing or wire protocol code. SBK’s normal writer/reader adapters provide the measurement timestamps. The synchronous driver path reports after the SDK call and response consumption complete. The asynchronous path reports from the tracked completion callback.
The SDK is intentionally pinned to 8.5.17 in
build.gradle. Later SDK behavior can add checksum
announcement headers that older ECS/ObjectScale and other S3-compatible
releases reject. An SDK upgrade is therefore a compatibility change, not a
routine dependency refresh.
| Class | Responsibility |
|---|---|
MinIO |
CLI/configuration binding, client construction, bucket setup, warm-up, catalog discovery, endpoint assignment, manifests, and lifecycle |
MinIOWriter |
PUT, update, copy, delete, tag mutation, and bucket mutation operations |
MinIOReader |
GET, Range GET, stat, tag read, LIST, bucket stat, and bucket list operations |
S3AsyncExecutor |
Per-worker and process-wide in-flight bounds; retains futures until their measurement callback completes |
S3MultipartUploader |
Bounded multipart waves, ordered completion, per-part retry, and abort after terminal failure |
S3ObjectCatalog |
Bounded existing-object inventory, reader partitioning, delete claims, and completed-PUT publication |
S3ObjectKey |
Sequential, hashed, random, filesystem-style, and distributed-partition key layouts |
S3ObjectSizeSelector |
Fixed, seeded uniform, deterministic sweep, and deterministic weighted object sizes |
S3OperationMix |
Exact repeating weighted operation cycles; no random selection noise |
S3RetryPolicy |
Optional retries for network I/O, HTTP 429, and HTTP 5xx responses |
S3DataGenerator |
Reproducible payload content, compressibility, and anti-dedup stamping |
S3PayloadPool |
Reuses arrays while regenerating object content; buffers are not reused until SDK completion |
S3EndpointMetrics |
Optional completed-operation, logical-byte, retry, and terminal-failure totals per configured endpoint |
MinIO.openStorage() performs startup work outside the measured operation
path:
-url list. A missing scheme
becomes http://.x-emc-namespace.-recreate true is explicit.Pure PUT, LIST, bucket-create, bucket-delete, bucket-stat, and bucket-list workloads do not require an existing-object catalog. GET, Range GET, stat, update, copy, delete, and object-tag operations do.
Catalog discovery is a recursive S3 LIST performed once at startup. It is
bounded by -catalog-max-objects. -object-file avoids that startup LIST by
loading key,size[,versionId] records from a local file. With
-partition-by-prefix true, distributed clients use a server-filterable
partition-N/ prefix; otherwise each client receives keys by stable Java hash.
Manifest parsing is intentionally strict: every non-comment line has exactly
two or three comma-separated fields, size is nonnegative, and object keys may
not contain commas.
One SBK record is one logical selected operation, including all configured retries. Multipart PUT is also one logical record even though it makes several S3 requests.
| Operation | Completion point | Bytes reported for that record |
|---|---|---|
| PUT/update | SDK upload completes | uploaded object bytes |
| Copy | server-side copy completes | source object logical bytes; bytes do not cross the load-generator network |
| Delete/tag set/tag delete/stat/tag get | SDK request completes | 0 |
| GET | complete response body is drained | response-body bytes |
| Range GET | requested response range is drained | returned range bytes |
| LIST | configured result entries/pages are consumed | 0; LIST transfers metadata, not object payload |
| Bucket create/delete/stat/list | SDK request completes | 0 |
LIST reports zero data bytes because Item.size() is the logical size of the
listed object, not bytes transferred by the metadata response. Judge LIST by
operations/sec and latency, and retain page size, entry limit, API version,
prefix, delimiter, and populated object count with the result.
SBK’s byte count for one operation is an int. Startup rejects GET/COPY
catalog entries and Range GET lengths above Integer.MAX_VALUE rather than
silently truncating their accounting.
-verify-read-size true checks GET and Range GET response lengths. It does not
compare response content with the original payload. -checksum asks the S3
service to validate the requested checksum on PUT; it is not a later GET
content comparison.
In synchronous mode, each SBK worker performs one SDK operation at a time. A comma-separated endpoint list assigns workers round-robin.
In asynchronous mode:
-async-depth SDK operations in flight;-async-max-inflight bounds all workers in the process;The conservative buffer estimate is:
writer payload bytes = min(process limit, writers * async depth) * largest object
reader buffer bytes = readers * async depth * 64 KiB
estimated bytes = writer payload bytes + reader buffer bytes
For synchronous execution, the estimate uses one object per writer and one
64 KiB buffer per reader. uniform, sweep, and weighted configurations use
their largest possible object in this calculation.
The OkHttp automatic total/per-host request limit is
max(64, workers * asyncDepth) unless explicitly overridden. Connection-pool
reuse avoids a TCP/TLS setup for every request.
-part-size selects multipart upload for eligible objects. The part size is
validated against S3’s 5 MiB through 5 GiB limits. With
-mpu-concurrent-parts 2 or greater, S3MultipartUploader:
The driver uses bounded views of the object payload instead of allocating a
new array for every part. Concurrent multipart cannot be combined with the
whole-object -checksum option.
For high-RTT ECS paths, first validate multipart with a small fixed
-records run. A timed run can have uploads in flight when the duration ends;
SBK will allow at most five seconds for all cleanup and will exit nonzero if
final results may be incomplete.
The default -retry-max-attempts 1 disables retries. When enabled, a retry is
allowed for network IOException, HTTP 429, and HTTP 5xx. The entire retry
sequence remains one latency sample, so retry delay and additional attempts
increase the reported operation latency. Fixed delay preserves compatibility;
exponential delay, an optional cap, and full jitter model production clients
and avoid coordinated retry waves. The timeout-classification change makes
active socket timeouts eligible for retry and prevents their classification as
clean shutdown. Only interruption caused by benchmark shutdown is treated as
clean shutdown.
When retries are enabled, a process-wide retry count is printed at shutdown
without adding bookkeeping to successful requests. -endpoint-metrics true
adds the opt-in per-endpoint success/byte/failure counters to the completion
path and also attributes retries by URL.
-endpoint-metrics true creates one counter group for each configured URL.
The shutdown summary contains completed logical operations/bytes, retry
attempts, and terminal failures. These are SDK completion counters rather than
PerL reporting-window counters: async completions at a timed boundary can make
them differ slightly from the final timed record count. A fixed-record
qualification should have exact completed-operation parity and zero failures.
Argument parsing resolves operation mixes, validates all booleans, headers, tags, endpoints, catalog limits, and finite target capacity before workers are created. A configured mix is authoritative; the single-operation option is only the fallback for an empty mix. Duplicate operation entries are rejected. This keeps configuration ambiguity and exhaustion checks out of measured requests.
The common successful PUT and full-GET paths retain their existing shape: one preselected operation, one prepared request, one SDK completion, and one SBK measurement. Range-offset selection runs only for the explicitly selected Range GET workload; LIST controls are prebuilt once per reader. Neither adds work to PUT or full GET. Retry scheduling runs only after a retryable failure. Detailed endpoint metrics remain opt-in because their completion counters do add bookkeeping.
If startup fails after any SDK client has been constructed, all constructed sync and async clients are closed and close failures are suppressed onto the primary startup exception. Warm-up cleanup follows the same rule: failure to remove a temporary object cannot hide the request failure that caused cleanup.
The JSON run manifest records the effective non-secret workload, sizing, integrity, retry, warm-up, HTTP, async, and partition configuration. It omits endpoint URLs, credentials, and extra-header values. This is a run-configuration record, not an object-result manifest.
For direct-IP ECS access, use the S3 data plane (9020 HTTP or 9021 HTTPS),
not the management UI/API, and inject the namespace header:
-url https://ecs.example.test:9021
-extra-headers x-emc-namespace=<namespace>
The header is installed in an OkHttp interceptor before the request proceeds. Credentials are the ECS Object User access key and secret key; management credentials are never used by this driver.
Worker close() drains tracked async futures. closeStorage() then removes
only buckets created by a bucket-create workload when requested, closes SDK
clients, and prints endpoint totals. Interrupt/rejected-executor failures caused
by normal timed teardown are treated as clean shutdown signals.
The enclosing SBK lifecycle has an absolute five-second cleanup deadline. If
workers, the SDK, the driver, PerL, or a logger cannot complete within that
deadline, the executable exits with failure because final results may be
incomplete. Never publish a run containing
BenchmarkCleanupTimeoutException.
The driver unit tests cover option validation, SDK argument construction, reader contracts, operation/size/key helpers, retry policy, payload reuse, async bounds, and multipart orchestration. Real-backend qualification remains required because SDK compatibility, TLS, namespace routing, permissions, and server behavior cannot be proven by unit tests.
Use:
./gradlew :drivers:minio:check :installDist
./build/install/sbk/bin/sbk -class minio -help
Then follow the fixed-record and timed stages in the ECS/OBS benchmark runbook.