Skip to content
Word Hunt SolverMy History

Validation

Word Hunt benchmarks: evidence before claims

Separate deterministic software checks from real-image recognition, game acceptance, and real-device performance.

Measured build evidence

The source release records exact command results, test cases, fixture labels, network checks, and simulated viewport screenshots in docs/VALIDATION.md. Published machine-readable results are available in the validation evidence record. A result only applies to the stated software revision and environment.

Source-release observations

CheckObserved resultConditions
Unit suite63 passed; 0 failed9 files · Cloud Node 24.19.0
Build / type check / lintPassedNext.js 16.4.0; 0 lint warnings
Browser suite18 passed; 0 failedCloud Chromium; desktop and simulated mobile; before OCR 1.1.0
Post-OCR targeted browser regressions8 passed; 0 failedLocal templates 1.1.0; explicit fallback suggestions; separate from the prior full suite
Public-source screenshot review check1 browser case; 3 imagesSame-case interface rerun; not a new blind evaluation

These are actual software checks. They do not estimate accuracy on game screenshots or performance on physical phones.

Solver correctness method

Tiny 2×2 and 3×3 boards are compared against an independent exhaustive path enumerator. Result paths are checked independently for range, adjacency, no reuse, token concatenation, minimum length, and dictionary membership. Larger-board tests specifically exercise tile indices beyond 31, cancellation, and budget-stopped partial results. These checks establish the tested algorithm behavior, not compatibility with a game’s private dictionary.

Screenshot recognition

Public-source GamePigeon UI images evaluated: 3, all 4×4. Two development images and one separate-source held-out image have unverified original device and installed game version. Local templates 1.1.0 produced 47 correct cells, one unknown and zero wrong cells; two whole boards were correct before editing. One natural correction recovered all three boards. Private consented corpus images remain zero. Generated PNG/JPEG/WebP regressions and corrected outcomes are counted separately. Read the actual OCR observations and protocol.

Game and device validation

Installed GamePigeon versions tested: 0. Real mobile devices tested: 0. Browser automation uses simulated desktop and mobile viewports on the cloud host. Read the game compatibility protocol before treating site scoring or dictionary output as a game guarantee.

Performance conditions matter

Cold measurements include resource download and initialization; hot measurements use an initialized dictionary or recognizer. The evidence record identifies the browser, viewport, fixture, word-list bytes, and timing boundary. A cloud Chromium timing is not an iPhone or Android benchmark, and a single observation is not a p95.

Practical details and worked examples

The documented benchmark methods are useful only when the measured task, environment, and sample are explicit. This release separates deterministic search correctness, synthetic image workflow checks, optional OCR initialization, and external-game compatibility. A successful browser workflow does not establish automatic real-image accuracy or physical-phone performance. Read the evidence labels before using a number to compare tools or deciding which part of the workflow needs improvement.

Identify the environment behind an observation

A reproducible run needs its software revision, browser, operating environment, viewport, dictionary, rules, and test input. A desktop Chromium run with a narrow viewport is a mobile layout simulation, not an iPhone or Android device measurement. Word Hunt solver benchmarks should identify which of those contexts was actually exercised. They should also distinguish an automated test passing from a user evaluation of readability, usability, or recognition accuracy.

Separate search correctness from resource agreement

Word Hunt solver benchmarks for correctness compare small-board results with an independent exhaustive approach and validate returned routes for range, adjacency, spelling, and no cell reuse. Large-grid tests address thirty-six and forty-nine positions, cancellation, revision races, and budget stopping. These checks concern the implemented engine under its selected dictionary. They do not show that an external game accepts the list, and a bounded set of alternate paths is not a count of every route.

Measure cold and warm work as different tasks

Cold Word Hunt solver benchmarks start with a first use, which may download a dictionary, construct a Trie, load a lazy API chunk, and initialize a WASM core or language model. A repeat use may reuse downloaded bytes and an initialized search resource. Report what was already loaded, what caching was enabled, and where timing starts and stops. A byte-size inventory is useful context but is not a cold-start duration. One quick warm run cannot represent every first visit.

Keep recognition and correction outcomes separate

Synthetic boards can exercise format decoding, crop controls, segmentation, candidates, and recovery. They cannot represent every game's glyphs, forwarded-image compression, or real-device themes. Word Hunt solver benchmarks must distinguish a transcribed board that was correct automatically from one made correct through manual review. All recognition results require review. Three source-published GamePigeon UI samples now provide scoped observations, with capture provenance and game versions unverified; their finite outcomes do not establish accuracy across devices or themes.

Reproduce evidence before generalizing from it

The repository contains test commands and recorded artifacts so a reviewer can rerun the relevant behavior. When results differ, first check the revision, settings, browser, and resources rather than averaging unrelated tasks. Public-source image observations are now recorded separately from synthetic checks. Consented current-version captures and physical devices remain listed gaps. A future comparison should use a predeclared sample protocol and report failures, incomplete searches, and corrected cases, not only the successful examples selected for a demonstration.

A practical sequence

  1. Name the exact task, resource versions, environment, and sample.
  2. Run correctness and lifecycle tests separately from recognition tests.
  3. Measure first initialization and repeat processing with defined timing boundaries.
  4. Publish observed outcomes together with failures, correction status, and untested conditions.

A worked example

Suppose a synthetic PNG decodes, the operator confirms its crop, recognition produces candidates, and two letters are corrected before solving. The completed board can demonstrate that the recovery flow works. It cannot count as an automatically correct screenshot. In a separate search test, matching an exhaustive small-board reference can support correctness under the selected words and rules. Neither observation estimates GamePigeon acceptance. A useful Word Hunt solver comparison preserves those distinctions instead of placing all three outcomes in one accuracy score.

Benchmark branches separate search correctness, image workflow, cold and warm timing, and external game acceptance.
A result supports its measured task and environment; it does not automatically transfer to another claim.

What to carry into your next attempt

Read Word Hunt solver benchmarks as scoped evidence. The most informative result explains what was measured, which inputs were used, and which claim remains unsupported. Use the OCR and dictionary-compatibility pages for their separate methods. Public-source image observations cover only their documented samples; physical-device performance, authenticated current-version recognition, and installed-game acceptance remain unverified.

Check the selected rules and custom points, word-list source and license, verification methods, and local data and network details when interpreting this example.

Updated October 10, 2026 · Independent Word Hunt Solver project