Extraction accuracy
The benchmark, including the uncomfortable fields.
On July 22, 2026, an evaluated TopLog extraction configuration read 3,512 of 3,528 compared hours cells correctly across 24 real logbook pages. Below are the field-level results and the method behind that number.
Hours cells
99.5%
3,512 of 3,528 correct
Count cells
98.0%
1,153 of 1,176 correct
Identity and text
93.5%
2,472 of 2,645 correct
Intermediate stops
84.6%
247 of 292 correct
Across hours, counts, identity, and text fields, the combined result was 97.1% (7,137 of 7,349). Intermediate stops are reported separately because a route can contain a variable-length list.
Independent review
The 24-page evaluation was independently reviewed.
Kris Mankala, CFI at Premiere Aviation, independently reviewed the benchmark. The affiliation identifies the reviewer and does not imply an organizational endorsement by Premiere Aviation.
Coverage
294 flights across 24 pages
The pages include different layouts, handwriting, aircraft, roles, route styles, hour columns, and training records. The frozen extraction output was compared against a human reference transcription.
Field-level results
Registration, SIC names, and intermediate stops were materially harder than hours. Publishing them beside the headline result makes the trade-off visible.
| Group | Field | Accuracy | Correct / compared |
|---|---|---|---|
| Flight identity | Date | 90.8% | 266 / 293 |
| Flight identity | Aircraft registration | 81.3% | 239 / 294 |
| Flight identity | Aircraft type | 98.3% | 289 / 294 |
| Flight identity | Departure | 95.6% | 281 / 294 |
| Flight identity | Arrival | 94.6% | 278 / 294 |
| Flight identity | Intermediate stops | 84.6% | 247 / 292 |
| Flight identity | Logged role | 96.6% | 284 / 294 |
| Flight identity | Engine type | 100.0% | 294 / 294 |
| Flight identity | PIC name | 99.7% | 293 / 294 |
| Flight identity | SIC name | 84.4% | 248 / 294 |
| Hours | Total | 99.0% | 291 / 294 |
| Hours | Day | 99.0% | 291 / 294 |
| Hours | Night | 99.7% | 293 / 294 |
| Hours | Cross-country | 100.0% | 294 / 294 |
| Hours | Ground training | 100.0% | 294 / 294 |
| Hours | Simulated instrument | 97.3% | 286 / 294 |
| Hours | FTD / simulator | 100.0% | 294 / 294 |
| Hours | Actual instrument | 100.0% | 294 / 294 |
| Hours | Dual given | 100.0% | 294 / 294 |
| Hours | Dual received | 100.0% | 294 / 294 |
| Hours | PICUS | 100.0% | 294 / 294 |
| Hours | Solo | 99.7% | 293 / 294 |
| Counts | Takeoffs / landings | 94.2% | 277 / 294 |
| Counts | Night landings | 99.7% | 293 / 294 |
| Counts | Instrument approaches | 98.3% | 289 / 294 |
| Counts | Holds | 100.0% | 294 / 294 |
How scoring works
The scoring rules are fixed before looking at the headline result.
- 1. Human reference. Each legible flight row is transcribed into structured reference data; cells ruled illegible are skipped.
- 2. Frozen output. The evaluated extraction configuration runs once against the same 24 pages and its output is saved before scoring.
- 3. Cell comparison. Numeric cells pass within ±0.05 hours. Text ignores case and surrounding whitespace; equivalent crew-name forms are matched semantically.
- 4. Separate failure types. Hours, counts, identity/text, and variable-length route stops are reported separately instead of hiding them inside one average.
What this benchmark does not prove
This is a bounded benchmark, not a guarantee for every logbook layout, photo condition, language, or handwriting style. Accuracy also varies by field, as the table shows. TopLog therefore presents extracted entries as drafts and the pilot must review the record before saving or relying on it.