# Current CJK evaluation protocol - frozen before output collection

Prepared 10 September 2026 for Plan044. Status: awaiting parent review of this protocol, rights, input hashes, rendered pages and ground truth. No OCR requests have been made. This is a small, single-provider descriptive evaluation, not a competitor benchmark.

## Frozen corpus and rights

Corpus root: `tmp/growth-044-infra/docs/quality-evidence/cjk-current/`. `input-manifest.json` records exact PDF, raster and ground-truth SHA256 values, generator hash, font hash and versions. Four original one-page documents: Japanese prose, Simplified Chinese prose, Japanese merged table and degraded English prose. All document text, table and raster artwork were newly composed for this evaluation, with no customer data or copied public document. Original material and generator are dedicated to CC0-1.0. The Noto Sans CJK Japanese font is redistributed with its SIL OFL 1.1 license in `assets/OFL.txt`; its glyphs are rasterized, with no font program embedded in the PDFs. Japanese font forms are used for both CJK cases; this is not a Chinese typography coverage test.

`generate.py` creates deterministic raster images with Pillow 12.3.0 and image-only PDFs with ReportLab 4.4.9, Python 3.12.14. Sources are 1240 x 1754 pixels at 150 dpi on 595.2 x 841.92 point pages. The degraded case is downsampled to 75 dpi, blurred with radius 0.75, rotated 0.7 degrees, reduced to intensity range 155..255, and given 7000 noise samples using seed 44. No request-side preprocessing is allowed. Exact bytes, not merely a rerun of the generator, define this corpus.

Every PDF has one page, empty pypdf extracted text and no text-showing operators. Poppler rendered every PDF for visual review. Review PNGs: `tmp/growth-044-infra/worker/node_modules/qa-044/{ja-prose,zh-prose,merged-table,degraded-scan}.png`. The executor checked every render for clipping, missing glyphs and agreement with the intended content; degradation is intentional. Parent must independently compare these pages with each case's `ground-truth.txt` before output collection. `merged-table/table-ground-truth.json` freezes 13 cells and three explicit spans (two rowspan=2, one colspan=2). Its text reading order is title, first header row, second header row, then body rows.

## Execution allocation and stopping rules

Unspent allocation: 12 conversion attempts maximum, three per case, one selected page `[0]` each. Customer payment budget $0, no purchase or paid identity. Nominal provider-processing ceiling $0.048 at the retained $0.004/page validation rate; actual provider charges are not observable and must not be invented. Only the existing public demo endpoint may be called, with application/json Base64 synthetic input, exactly `{"input":{"pdf_base64":"<exact input.pdf bytes as Base64>","max_pages":1,"include_raw":true}}`; never canonical sample cache URLs. No credentials are required or requested.

Order: ja-prose 1..3, zh-prose 1..3, merged-table 1..3, degraded-scan 1..3. Identical bytes, selected page and transport for repetitions. Start requests at least 25 seconds apart (within the documented three/minute demo rate), sequentially with no hidden retries and no redirects that replay a request. Use a 120-second timeout. Every dispatched conversion consumes one attempt, including timeout, 429 or any other error. Never replace failures. Stop at the cap or public-demo unavailability; preserve attempted failures and explicitly mark remaining attempts unattempted. Root review is mandatory before the first request.

## Captures and metrics

Record UTC start, wall-clock elapsed duration measured around the whole public request, HTTP status, raw-PDF/Base64 transport, exact source hash, selected pages, actual x-ocr-model and x-ocr-provider headers (null if absent), `response.raw.model`, cache/sample headers and complete sanitized public result. The reviewed worker strips internal model/provider/cache headers, so pin the observed model only from `response.raw.model` returned with `include_raw=true`; missing raw.model remains unknown. Never substitute a configured model for an observed one. Sanitize only request identifiers, internal/private headers and sensitive metadata; never edit the returned Markdown. Record missing model/cache headers as null/unknown. `worker/src/sample-cache.ts` at infra bd06d7041e0286807fbed8d81c03c9c0d37e1118 disables the prepared sample cache when includeRaw is true or the source hash differs from the canonical invoice; these original inputs satisfy both, and no canonical pdf_url is supplied. This establishes service sample-cache ineligibility from reviewed code and bytes, not a live cache header. Provider-internal caching remains unknown, including for identical repetitions. Provider timing headers are not end-to-end timing. Expected current model is mistral-ocr-4-1; a different or missing actual model cannot be promoted to a passing current-model report.

For each case, designate repetition 1 as the primary output before collection. Compute its text CER using the existing publication verifier's exact normalization: strip HTML tags, heading markers, GFM separators/pipes and demo watermark; NFKC; remove whitespace; Unicode code-point Levenshtein distance divided by normalized ground-truth length. Do not select the best repetition. Retain all other outputs. Table cells and explicit HTML spans are separate exact checks, not CER; no percentage layout score or claim that text accuracy establishes layout accuracy.

Nearest-rank median and p95 require all three planned successful comparable repetitions that are ineligible for the service sample cache and passing frozen expected-text/table checks. With n=3 the p95 is only the maximum of three observations, not a stable tail estimate. This small corpus cannot establish language-wide quality or superiority. Do not publish a success-only timing subset. If a case has any failed attempt, changed/missing model, unresolved service sample-cache eligibility or failed quality check, preserve the entire case as failed/unmeasured with published metrics null. Do not weaken the existing verifier. Failed captures may use a separate capture ledger rather than the successful publication schema. Every published numeric metric must recompute from retained artifacts. Keep invoice and July evidence unchanged.

## Required checkpoint

Parent review response must be retained before collection, identifying accepted corpus hashes and any corrections. Any change to input bytes, ground truth, settings or scoring after review requires a new review before outputs. Execution status and actual budget use will be recorded separately so this pre-output protocol stays frozen.

Parent first review accepted all four input hashes, rendered pages and ground truth unchanged; requested include_raw=true and source-based service-cache eligibility instead of requiring stripped internal headers. The protocol incorporates that correction before requests. Exact collector: `tmp/growth-044-infra/docs/quality-evidence/cjk-current/collect.py`; it consumes one ledger slot before each dispatch, disables redirects/retries, and requires retained review hashes. A second parent review of these options and the collector remains pending.
