Skip to main content

Extraction accuracy

The benchmark, including the uncomfortable fields.

On July 22, 2026, an evaluated TopLog extraction configuration read 3,512 of 3,528 compared hours cells correctly across 24 real logbook pages. Below are the field-level results and the method behind that number.

Hours cells

99.5%

3,512 of 3,528 correct

Count cells

98.0%

1,153 of 1,176 correct

Identity and text

93.5%

2,472 of 2,645 correct

Intermediate stops

84.6%

247 of 292 correct

Across hours, counts, identity, and text fields, the combined result was 97.1% (7,137 of 7,349). Intermediate stops are reported separately because a route can contain a variable-length list.

Independent review

The 24-page evaluation was independently reviewed.

Kris Mankala, CFI at Premiere Aviation, independently reviewed the benchmark. The affiliation identifies the reviewer and does not imply an organizational endorsement by Premiere Aviation.

Coverage

294 flights across 24 pages

The pages include different layouts, handwriting, aircraft, roles, route styles, hour columns, and training records. The frozen extraction output was compared against a human reference transcription.

Field-level results

Registration, SIC names, and intermediate stops were materially harder than hours. Publishing them beside the headline result makes the trade-off visible.

Accuracy by extracted logbook field
GroupFieldAccuracyCorrect / compared
Flight identityDate90.8%266 / 293
Flight identityAircraft registration81.3%239 / 294
Flight identityAircraft type98.3%289 / 294
Flight identityDeparture95.6%281 / 294
Flight identityArrival94.6%278 / 294
Flight identityIntermediate stops84.6%247 / 292
Flight identityLogged role96.6%284 / 294
Flight identityEngine type100.0%294 / 294
Flight identityPIC name99.7%293 / 294
Flight identitySIC name84.4%248 / 294
HoursTotal99.0%291 / 294
HoursDay99.0%291 / 294
HoursNight99.7%293 / 294
HoursCross-country100.0%294 / 294
HoursGround training100.0%294 / 294
HoursSimulated instrument97.3%286 / 294
HoursFTD / simulator100.0%294 / 294
HoursActual instrument100.0%294 / 294
HoursDual given100.0%294 / 294
HoursDual received100.0%294 / 294
HoursPICUS100.0%294 / 294
HoursSolo99.7%293 / 294
CountsTakeoffs / landings94.2%277 / 294
CountsNight landings99.7%293 / 294
CountsInstrument approaches98.3%289 / 294
CountsHolds100.0%294 / 294

How scoring works

The scoring rules are fixed before looking at the headline result.

  1. 1. Human reference. Each legible flight row is transcribed into structured reference data; cells ruled illegible are skipped.
  2. 2. Frozen output. The evaluated extraction configuration runs once against the same 24 pages and its output is saved before scoring.
  3. 3. Cell comparison. Numeric cells pass within ±0.05 hours. Text ignores case and surrounding whitespace; equivalent crew-name forms are matched semantically.
  4. 4. Separate failure types. Hours, counts, identity/text, and variable-length route stops are reported separately instead of hiding them inside one average.

What this benchmark does not prove

This is a bounded benchmark, not a guarantee for every logbook layout, photo condition, language, or handwriting style. Accuracy also varies by field, as the table shows. TopLog therefore presents extracted entries as drafts and the pilot must review the record before saving or relying on it.