phypa.org · Guides

What to record before comparing EEG benchmark results

Published October 5, 2026 · Independent editorial guide

An EEG classification result is difficult to interpret without the procedure that produced it. A headline accuracy can hide differences in participants, recording sessions, task labels, preprocessing or evaluation splits. Before comparing two approaches, prepare a short benchmark record that lets another researcher identify those differences.

The PhyPA benchmarking section provides historical context for the site's interest in physiological parameters and human-machine interaction. This independent methodological note reports no new experiment and makes no clinical-performance claim.

Identify the exact data and labels

Record the dataset name, version, files used and any exclusion rule. The PhysioNet EEG Motor Movement/Imagery dataset supplies recording and task documentation alongside the data. Read that documentation before assigning labels; movement and imagined-movement recordings should not be treated as interchangeable by assumption. Cite the dataset and its requested references when reporting work that uses it.

Define the comparison before developing the method

State the evaluation question and keep development decisions separate from held-out evaluation. Record whether the comparison concerns new trials, sessions or participants. These answer different generalization questions. Document how preprocessing is fitted and which data it sees, so an evaluation does not quietly benefit from information intended to be held out.

Write the split specification in a reproducible form and retain the random seed where applicable. If an exclusion rule changes after results are viewed, report the change rather than describing the new analysis as the original plan.

Keep an environment record

List operating system, language runtime, library versions, configuration and relevant hardware. Preserve a small test that exercises the intended workflow. The used-workstation assessment guide provides a useful equipment checklist when an older machine is part of a research environment. A historic hardware specification does not establish compatibility with a current analysis tool.

Report uncertainty and limitations

Present the metric definition, evaluation unit and variation across the relevant observations. Do not describe a change in a single summary score as universal improvement. Keep error cases and departures from the planned procedure visible. The guide to reliable observations and feedback offers a broader account of separating a measurement from the explanation attached to it.

A benchmark record can be short: data, labels, split, processing, environment, metric and limitations. The important property is that every comparison refers to the same documented question. If a result cannot be reconstructed from those notes, improving the record is a worthwhile next step before claiming a stronger method.