dantinox.benchmarking

The benchmarking module is a plugin framework: tasks are independent classes, the suite orchestrates them, and results aggregate into a structured report.


Suite orchestrator

class dantinox.benchmarking.suite.BenchmarkSuite(tasks: list[BenchmarkTask], config: BenchmarkConfig | None = None)[source]

Bases: object

Runs a collection of BenchmarkTask instances against a model.

The suite is the single entry-point for evaluation. It owns no model logic — it delegates entirely to each task’s run() method.

Quick-start:

from dantinox.benchmarking import BenchmarkSuite

# Default suite: Throughput + Latency + Perplexity
report = BenchmarkSuite.default().run(paradigm, model)
print(report.summary())
report.save("results.csv")

Custom suite:

suite = BenchmarkSuite(
    tasks=[ThroughputTask(), PerplexityTask("data/val.txt")],
    config=BenchmarkConfig(seq_lens=[128, 512], n_measure=50),
)
report = suite.run(paradigm, model)
run(paradigm: Any, model: Any, *, save_csv: str | None = None) → SuiteReport[source]

Run all tasks sequentially and return a SuiteReport.

Parameters:
  • paradigm – Any Paradigm instance.

  • model – The NNX model returned by paradigm.build_model().

  • save_csv – If provided, write the report to this CSV path.

Returns:

A SuiteReport aggregating every task’s outcome.

classmethod default(config: BenchmarkConfig | None = None) → BenchmarkSuite[source]

Return a suite with the three standard tasks: Throughput + Latency + Perplexity.

Perplexity is only meaningful when a data source is attached at runtime; if no data is available the task will log a warning and return NaN.

classmethod throughput_only(config: BenchmarkConfig | None = None) → BenchmarkSuite[source]

Return a minimal suite for quick hardware throughput checks.


Plugin base class

class dantinox.benchmarking.base.BenchmarkTask[source]

Bases: ABC

Abstract base for all benchmark tasks.

Implementing a new task requires only one override:

class MyTask(BenchmarkTask):
    name = "my_task"

    def run(self, paradigm, model, config, rng):
        score = evaluate_something(model)
        return BenchmarkResult(task=self.name, metrics={"score": score})

The task is then plug-and-play with any BenchmarkSuite:

suite = BenchmarkSuite([MyTask(), ThroughputTask()])
name: ClassVar[str]
abstractmethod run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Any) → BenchmarkResult[source]

Execute the task and return a BenchmarkResult.

Parameters:
  • paradigm – The Paradigm wrapping the model.

  • model – The NNX model object (Transformer or FlowMatchingTransformer).

  • config – Suite-level benchmark configuration.

  • rng – JAX random key.


Result types

class dantinox.benchmarking.base.BenchmarkResult(task: str, metrics: dict[str, float], profiling: ~dantinox.profiling.tracker.ProfilingResult | None = None, flops: ~dantinox.profiling.counter.FLOPsBreakdown | None = None, meta: dict[str, ~typing.Any] = <factory>)[source]

Bases: object

Outcome of a single BenchmarkTask run.

metrics holds the task-specific scalars (e.g. perplexity, accuracy, tok/s). profiling and flops are populated automatically by tasks that measure hardware efficiency.

task: str
metrics: dict[str, float]
profiling: ProfilingResult | None = None
flops: FLOPsBreakdown | None = None
meta: dict[str, Any]
to_dict() → dict[str, Any][source]
class dantinox.benchmarking.base.SuiteReport(results: list[~dantinox.benchmarking.base.BenchmarkResult], model_meta: dict[str, ~typing.Any] = <factory>, total_time_s: float = 0.0)[source]

Bases: object

Aggregated output of a BenchmarkSuite run.

Attributes

results : One BenchmarkResult per task that was executed. model_meta : Static model info (n_params, paradigm name, config repr…). total_time_s : Wall-clock seconds for the entire suite.

results: list[BenchmarkResult]
model_meta: dict[str, Any]
total_time_s: float = 0.0
to_dataframe() → Any[source]

Return a pandas.DataFrame with one row per task.

save(path: str) → None[source]

Save the suite results to path (CSV).

summary() → str[source]
class dantinox.benchmarking.base.BenchmarkConfig(seq_lens: list[int] = <factory>, batch_sizes: list[int] = <factory>, n_warmup: int = 5, n_measure: int = 20, eval_batches: int = 50, eval_seq_len: int = 256, eval_batch_size: int = 4, output_dir: str = 'benchmark_results', seed: int = 0)[source]

Bases: object

Hardware and evaluation settings for a benchmark suite.

All fields have sensible defaults so a zero-config run is possible:

suite.run(model, paradigm)  # uses BenchmarkConfig defaults
seq_lens: list[int]
batch_sizes: list[int]
n_warmup: int = 5
n_measure: int = 20
eval_batches: int = 50
eval_seq_len: int = 256
eval_batch_size: int = 4
output_dir: str = 'benchmark_results'
seed: int = 0
to_dict() → dict[str, Any][source]
classmethod from_dict(d: dict[str, Any]) → BenchmarkConfig[source]
classmethod from_yaml(path: str) → BenchmarkConfig[source]
save_yaml(path: str) → None[source]

Built-in tasks

class dantinox.benchmarking.tasks.throughput.ThroughputTask[source]

Bases: BenchmarkTask

Measure decode throughput (tokens/s) across sequence lengths and batch sizes.

Uses the profiling LatencyTracker to record wall-clock latency and derive tokens/s. Both a seq-len sweep (batch=1) and a batch-size sweep (fixed seq_len) are reported.

Metrics produced

tps_seq{L} : tokens/s at sequence length L, batch=1. tps_bs{B} : tokens/s at batch size B, seq_len=config.seq_lens[0]. peak_tps : highest measured throughput across all configurations.

name: ClassVar[str] = 'throughput'
run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) → BenchmarkResult[source]

Execute the task and return a BenchmarkResult.

Parameters:
  • paradigm – The Paradigm wrapping the model.

  • model – The NNX model object (Transformer or FlowMatchingTransformer).

  • config – Suite-level benchmark configuration.

  • rng – JAX random key.

class dantinox.benchmarking.tasks.latency.LatencyTask[source]

Bases: BenchmarkTask

Measure prefill and (for AR models) decode latency.

Prefill is defined as a single full-sequence forward pass. Decode is defined as a single next-token generation step, applicable only to autoregressive paradigms.

Metrics produced

prefill_mean_ms : mean prefill latency in ms (prompt_len = seq_lens[-1]). prefill_p99_ms : 99th-percentile prefill latency. decode_mean_ms : mean single-step decode latency in ms (AR only). decode_p99_ms : 99th-percentile decode latency (AR only). decode_tps : 1 / decode_mean_s (AR only).

name: ClassVar[str] = 'latency'
run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) → BenchmarkResult[source]

Execute the task and return a BenchmarkResult.

Parameters:
  • paradigm – The Paradigm wrapping the model.

  • model – The NNX model object (Transformer or FlowMatchingTransformer).

  • config – Suite-level benchmark configuration.

  • rng – JAX random key.

class dantinox.benchmarking.tasks.perplexity.PerplexityTask(data_source: str | None = None)[source]

Bases: BenchmarkTask

Compute perplexity on a held-out text corpus via the paradigm’s loss.

Works with any paradigm that implements loss_fn(model, batch, rng) returning a cross-entropy-based scalar. For AR models this is the standard next-token perplexity; for diffusion models it is the ELBO-weighted surrogate (still a valid quality proxy).

Parameters:

data_source – Path to a plain-text file used as the evaluation corpus. If None, the task skips and returns NaN.

Metrics produced

perplexity : exp(mean cross-entropy loss over eval batches). eval_loss : raw mean cross-entropy.

name: ClassVar[str] = 'perplexity'
run(paradigm: Any, model: Any, config: BenchmarkConfig, rng: Array) → BenchmarkResult[source]

Execute the task and return a BenchmarkResult.

Parameters:
  • paradigm – The Paradigm wrapping the model.

  • model – The NNX model object (Transformer or FlowMatchingTransformer).

  • config – Suite-level benchmark configuration.

  • rng – JAX random key.


Quick reference

from dantinox.benchmarking import BenchmarkSuite, BenchmarkConfig

# Default suite (Throughput + Latency + Perplexity)
report = BenchmarkSuite.default().run(paradigm, model)

# Custom suite
from dantinox.benchmarking.tasks.perplexity import PerplexityTask
suite  = BenchmarkSuite(
    tasks=[PerplexityTask("data/val.txt")],
    config=BenchmarkConfig(eval_batches=100, eval_seq_len=512),
)
report = suite.run(paradigm, model, save_csv="results.csv")
print(report.summary())
df = report.to_dataframe()

Metrics produced by built-in tasks

Task

Metric key

Description

ThroughputTask

tps_seq{L}

Tokens/s at sequence length L, batch=1

ThroughputTask

tps_bs{B}

Tokens/s at batch size B

ThroughputTask

peak_tps

Maximum observed tokens/s

LatencyTask

prefill_mean_ms

Mean prefill latency

LatencyTask

prefill_p99_ms

99th-percentile prefill latency

LatencyTask

decode_mean_ms

Mean single-step decode latency (AR only)

LatencyTask

decode_tps

Decode throughput (AR only)

PerplexityTask

perplexity

exp(mean_ce_loss)

PerplexityTask

eval_loss

Mean cross-entropy loss