Korean models & evaluation
Evaluation designed for Korean instructions, long coding sessions, local documents, and repository conventions.
Patty reports do more than display results. They explain why the question matters, the conditions behind the measurement, how a decision changes, and what the evidence cannot establish.
We study the complete deployable path rather than an isolated model score.
Evaluation designed for Korean instructions, long coding sessions, local documents, and repository conventions.
Measurements across distributed serving, scheduling, KV cache, and resource boundaries.
Finished-path analysis connecting observation, planning, approval, tool execution, and recovery.
Validation that identity, delegation, data boundaries, and execution receipts operate as real controls.
Inspect the measurement unit and exclusion boundary before the size of a result.
The design or operating decision the report addresses
Hardware, build, data, participants, and observation period
What was observed and how it changes a decision
What cannot be generalized or has not yet been verified
A measurement of the complete observe→classify→park→approve→dispatch path when an AI agent is embedded in the browser engine. Semantic snapshots remained around 14–16ms in the tested fixtures, and the governed engine cycle completed in under 25ms.
Design notes for an end-to-end corpus stratified across eight workload classes, a 50-instance SWE-bench Verified subset, and a context-retention A/B harness, including anti-guess tasks and scorer validation.
Reports evolve with releases, remeasurement, and environmental change. Modification dates and measurement boundaries remain explicit; earlier results are not silently generalized.
Instead of copying a public number, design an evaluation around your hardware, data, network, and approval requirements with Patty.