dantinox.benchmarking
The benchmarking module is a plugin framework: tasks are independent classes, the suite orchestrates them, and results aggregate into a structured report.
Suite orchestrator
- class dantinox.benchmarking.suite.BenchmarkSuite(tasks: list[BenchmarkTask], config: BenchmarkConfig | None = None)[source]
Bases:
objectRuns a collection of
BenchmarkTaskinstances against a model.The suite is the single entry-point for evaluation. It owns no model logic — it delegates entirely to each task’s
run()method.Quick-start:
from dantinox.benchmarking import BenchmarkSuite # Default suite: Throughput + Latency + Perplexity report = BenchmarkSuite.default().run(paradigm, model) print(report.summary()) report.save("results.csv")
Custom suite:
suite = BenchmarkSuite( tasks=[ThroughputTask(), PerplexityTask("data/val.txt")], config=BenchmarkConfig(seq_lens=[128, 512], n_measure=50), ) report = suite.run(paradigm, model)
- run(paradigm: Any, model: Any, *, save_csv: str | None = None) SuiteReport[source]
Run all tasks sequentially and return a
SuiteReport.- Parameters:
paradigm – Any
Paradigminstance.model – The NNX model returned by
paradigm.build_model().save_csv – If provided, write the report to this CSV path.
- Returns:
A
SuiteReportaggregating every task’s outcome.
- classmethod default(config: BenchmarkConfig | None = None) BenchmarkSuite[source]
Return a suite with the three standard tasks: Throughput + Latency + Perplexity.
Perplexity is only meaningful when a data source is attached at runtime; if no data is available the task will log a warning and return NaN.
- classmethod throughput_only(config: BenchmarkConfig | None = None) BenchmarkSuite[source]
Return a minimal suite for quick hardware throughput checks.
Plugin base class
- class dantinox.benchmarking.base.BenchmarkTask[source]
Bases:
ABCAbstract base for all benchmark tasks.
Implementing a new task requires only one override:
class MyTask(BenchmarkTask): name = "my_task" def run(self, paradigm, model, config, rng): score = evaluate_something(model) return BenchmarkResult(task=self.name, metrics={"score": score})
The task is then plug-and-play with any
BenchmarkSuite:suite = BenchmarkSuite([MyTask(), ThroughputTask()])
- abstractmethod run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Any) BenchmarkResult[source]
Execute the task and return a
BenchmarkResult.- Parameters:
paradigm – The
Paradigmwrapping the model.model – The NNX model object (Transformer or FlowMatchingTransformer).
config – Suite-level benchmark configuration.
rng – JAX random key.
Result types
- class dantinox.benchmarking.base.BenchmarkResult(task: str, metrics: dict[str, float], profiling: ~dantinox.profiling.tracker.ProfilingResult | None = None, flops: ~dantinox.profiling.counter.FLOPsBreakdown | None = None, meta: dict[str, ~typing.Any] = <factory>)[source]
Bases:
objectOutcome of a single
BenchmarkTaskrun.metrics holds the task-specific scalars (e.g. perplexity, accuracy, tok/s). profiling and flops are populated automatically by tasks that measure hardware efficiency.
- profiling: ProfilingResult | None = None
- flops: FLOPsBreakdown | None = None
- class dantinox.benchmarking.base.SuiteReport(results: list[~dantinox.benchmarking.base.BenchmarkResult], model_meta: dict[str, ~typing.Any] = <factory>, total_time_s: float = 0.0)[source]
Bases:
objectAggregated output of a
BenchmarkSuiterun.Attributes
results : One
BenchmarkResultper task that was executed. model_meta : Static model info (n_params, paradigm name, config repr…). total_time_s : Wall-clock seconds for the entire suite.- results: list[BenchmarkResult]
- class dantinox.benchmarking.base.BenchmarkConfig(seq_lens: list[int] = <factory>, batch_sizes: list[int] = <factory>, n_warmup: int = 5, n_measure: int = 20, eval_batches: int = 50, eval_seq_len: int = 256, eval_batch_size: int = 4, output_dir: str = 'benchmark_results', seed: int = 0)[source]
Bases:
objectHardware and evaluation settings for a benchmark suite.
All fields have sensible defaults so a zero-config run is possible:
suite.run(model, paradigm) # uses BenchmarkConfig defaults
- classmethod from_yaml(path: str) BenchmarkConfig[source]
Built-in tasks
- class dantinox.benchmarking.tasks.throughput.ThroughputTask[source]
Bases:
BenchmarkTaskMeasure decode throughput (tokens/s) across sequence lengths and batch sizes.
Uses the profiling
LatencyTrackerto record wall-clock latency and derive tokens/s. Both a seq-len sweep (batch=1) and a batch-size sweep (fixed seq_len) are reported.Metrics produced
tps_seq{L}: tokens/s at sequence length L, batch=1.tps_bs{B}: tokens/s at batch size B, seq_len=config.seq_lens[0].peak_tps: highest measured throughput across all configurations.- run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) BenchmarkResult[source]
Execute the task and return a
BenchmarkResult.- Parameters:
paradigm – The
Paradigmwrapping the model.model – The NNX model object (Transformer or FlowMatchingTransformer).
config – Suite-level benchmark configuration.
rng – JAX random key.
- class dantinox.benchmarking.tasks.latency.LatencyTask[source]
Bases:
BenchmarkTaskMeasure prefill and (for AR models) decode latency.
Prefill is defined as a single full-sequence forward pass. Decode is defined as a single next-token generation step, applicable only to autoregressive paradigms.
Metrics produced
prefill_mean_ms: mean prefill latency in ms (prompt_len = seq_lens[-1]).prefill_p99_ms: 99th-percentile prefill latency.decode_mean_ms: mean single-step decode latency in ms (AR only).decode_p99_ms: 99th-percentile decode latency (AR only).decode_tps: 1 / decode_mean_s (AR only).- run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) BenchmarkResult[source]
Execute the task and return a
BenchmarkResult.- Parameters:
paradigm – The
Paradigmwrapping the model.model – The NNX model object (Transformer or FlowMatchingTransformer).
config – Suite-level benchmark configuration.
rng – JAX random key.
- class dantinox.benchmarking.tasks.perplexity.PerplexityTask(data_source: str | None = None)[source]
Bases:
BenchmarkTaskCompute perplexity on a held-out text corpus via the paradigm’s loss.
Works with any paradigm that implements
loss_fn(model, batch, rng)returning a cross-entropy-based scalar. For AR models this is the standard next-token perplexity; for diffusion models it is the ELBO-weighted surrogate (still a valid quality proxy).- Parameters:
data_source – Path to a plain-text file used as the evaluation corpus. If
None, the task skips and returnsNaN.
Metrics produced
perplexity: exp(mean cross-entropy loss over eval batches).eval_loss: raw mean cross-entropy.- run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) BenchmarkResult[source]
Execute the task and return a
BenchmarkResult.- Parameters:
paradigm – The
Paradigmwrapping the model.model – The NNX model object (Transformer or FlowMatchingTransformer).
config – Suite-level benchmark configuration.
rng – JAX random key.
Quick reference
from dantinox.benchmarking import BenchmarkSuite, BenchmarkConfig
# Default suite (Throughput + Latency + Perplexity)
report = BenchmarkSuite.default().run(paradigm, model)
# Custom suite
from dantinox.benchmarking.tasks.perplexity import PerplexityTask
suite = BenchmarkSuite(
tasks=[PerplexityTask("data/val.txt")],
config=BenchmarkConfig(eval_batches=100, eval_seq_len=512),
)
report = suite.run(paradigm, model, save_csv="results.csv")
print(report.summary())
df = report.to_dataframe()
Metrics produced by built-in tasks
Task |
Metric key |
Description |
|---|---|---|
|
|
Tokens/s at sequence length L, batch=1 |
|
|
Tokens/s at batch size B |
|
|
Maximum observed tokens/s |
|
|
Mean prefill latency |
|
|
99th-percentile prefill latency |
|
|
Mean single-step decode latency (AR only) |
|
|
Decode throughput (AR only) |
|
|
|
|
|
Mean cross-entropy loss |