OCR benchmark
Word Hunt screenshot benchmarks: method and limits
Three public-source GamePigeon UI images were evaluated. Their device and installed game version are unverified; this small convenience sample does not establish general screenshot accuracy.
Actual sample scope
The frozen manifest contains 2 gameplay images and 1 instruction image, all 4×4. Two development images informed recognition fixes; one image from a separate source was held out from those changes. The later review-interface check reused the same three images and is not a second blind evaluation. Private consented corpus images and physical mobile devices tested: zero. Original image pixels are not redistributed in this source release.
Observed automatic outcomes
| Stage | Located boards | Correct cells | Unknown cells | Wrong nonempty cells | Entirely correct boards |
|---|---|---|---|---|---|
| Before development fixes | 2/3 | 29/48 | 17 | 2 | 1/3 |
| Local templates 1.1.0 | 3/3 | 47/48 | 1 | 0 | 2/3 |
Labels and board boundaries were frozen before recognition. The new result includes development reuse and a single separate-source held-out board. It is an observed count for these images, not a representative accuracy estimate. The remaining unknown letter was covered by a selected-tile overlay.
Fallback suggestions require a decision
Raw local Tesseract proposals produced 29 correct, 16 unknown, and 3 wrong cells. An earlier automatic merge reduced correct cells to 45, so that policy was replaced. The current interface preserves primary letters, checked flags, manual corrections and undo history; proposals appear separately and only an explicit per-cell click applies one. A wrong proposal was deliberately applied and undone during the interface check. Fallback availability does not mean its proposal is more accurate.
Human recovery is separate
One natural correction was needed after primary recognition. All 3 boards were reviewed, confirmed and solved. That recovery result is separate from the 2 unedited whole-board successes. Deliberate interface-test edits are excluded from the natural correction count. Correction-time accuracy and population-wide rates were not measured.
Reproduce and extend the method
Inspect the measured evidence record for source pages, image fingerprints, frozen labels, algorithm hashes and report paths. Extend evaluation with consented images from different devices, themes and sizes; freeze truth before tuning, split by source board, and retain failures. Measure localization, cell transcription, complete boards and manual recovery separately. Generated PNG/JPEG/WebP tests establish regression behavior rather than real-game accuracy.
Cold processing and limits
Cloud Chromium resource and initialization observations are recorded separately from warmed dictionary searches. Repeated fallback uses cached assets but creates a new OCR engine. No physical-phone timing, HEIC support, severe perspective correction, or broad unknown-font accuracy is established. Try the local screenshot workflow; every tile still requires review.
Practical details and worked examples
Automatic recognition quality needs more than a successful demonstration image. A Word Hunt screenshot accuracy report separates board localization, cell transcription, review coverage, and the time spent correcting errors. The current evidence includes synthetic fixtures, browser workflows, and three source-published GamePigeon UI images with independently frozen labels. Capture provenance and game versions remain unverified. The recorded observations distinguish two development images from one held-out image; physical-device samples are still absent, and a tiny public-source sample cannot establish a general real-game rate.
Identify the sample before calculating a rate
Label each fixture by format, board size, glyph style, resolution, and whether its crop was supplied or chosen by a user. The current synthetic sample workflow covers twelve format-and-size combinations with manual crop confirmation. That exercises an important path through the interface, but it is not an unassisted localization test on a game conversation. Word Hunt screenshot accuracy should always name the denominator and exclude claims that the sample did not test.
Keep automatic transcription and corrected boards separate
A Word Hunt screenshot accuracy evaluation compares the recognizer's original character output with a known transcription. Whole-board correctness requires every tile to match. A manual correction changes the outcome to a corrected success, even if only one cell needed editing. Record crop assistance separately as well. A finished Word Hunt solver screenshot may be fully usable after review, while the automatic whole-board result for the same attempt was wrong or unknown. Both observations can be useful when reported honestly.
Do not interpret similarity as calibrated confidence
The local template matcher produces similarity scores and candidate letters. Those numbers are review aids, not measured probabilities of correctness. A high-scoring wrong letter is still wrong, and an unknown tile is not evidence that the original board lacked a letter. The interface requires checking all cells, including apparently strong suggestions. Any future automatic-acceptance policy would need separate blind evidence about false approvals and coverage before bypassing that review gate.
Define the untested styles and devices
Real-device Word Hunt screenshot accuracy is unmeasured: generic runtime font templates do not establish compatibility with a particular GamePigeon theme or application version. Current missing conditions include consented and authenticated current-version captures, physical iOS Safari and Android Chrome, broad forwarded-image compression samples, and angled camera photographs. Publicly sourced images expand the observed inputs without resolving those device and provenance gaps. Limited rotation correction is implemented; perspective correction is not validated. A synthetic narrow viewport screenshot does not measure a phone's decode time, memory pressure, thermal state, touch behavior, or recognition performance.
Plan an authorized blind evaluation
Obtain permission for sample use, remove conversation content, and create a held-out set whose letters are transcribed independently. Freeze the resource version before running it. Record localization failures, unknown cells, incorrect confident suggestions, unassisted board correctness, corrected completion, and correction time. Measure first download and repeat recognition separately. Word Hunt screenshot accuracy becomes meaningful when failures and assistance are visible rather than hidden behind a single polished example.
A practical sequence
- Label every sample and record whether crop assistance was provided.
- Capture original recognized letters before manual changes.
- Score characters and whole boards separately, then record corrected success.
- Publish authorized sample scope, resource versions, failures, and timing conditions.
A worked example
For an illustrative calculation, suppose a 6×6 synthetic WebP contains thirty-six known letters. The user chooses the crop; recognition returns thirty-four correct letters, one wrong suggestion, and one unknown. After review, both cells are repaired and the board solves. The corrected workflow succeeded, but the automatic whole-board attempt did not. Its manually supplied crop also cannot count as automatic localization. These are different measurements for one Word Hunt solver screenshot, and retaining all of them makes the result interpretable.
What to carry into your next attempt
Use Word Hunt screenshot accuracy to describe an explicitly labeled sample and stage, not a universal capability. The current product provides review and manual recovery because a small public-source sample does not establish general automatic reliability. Consult the recorded observations above for original output, unknown cells, and corrected completion. Future evaluations should preserve those distinctions and publish their sample split, failures, correction effort, and cold-versus-warm timing before making stronger claims.
Check the selected rules and custom points, word-list source and license, verification methods, and local data and network details when interpreting this example.
Updated October 10, 2026 · Independent Word Hunt Solver project