Armenian Corpus of Printed Periodicals โ OCR on 197 scanned pages of historical Armenian newspapers (1962โ1985). Models are ranked by CER (character error rate), lower is better. ๐ Dataset
Search
Model type
Select columns to show
Filter
Structure-awareStructure-independentยท shorter bar = lower CER = better
About the benchmark
The Armenian Corpus of Printed Periodicals (ACoPPer) is an expanding digital archive of written texts. Its test split contains 197 scanned pages from Armenian newspapers and periodicals of 1962โ1985, with detailed layout annotations (38 region categories) and ground-truth transcriptions.
Settings
Structure-aware โ the model outputs one complete text line per box; its own line segmentation is scored as-is.
Structure-independent โ the model outputs word boxes, which are grouped into lines by a fixed heuristic before scoring.
Filters โ exclude non-Armenian text, graphics & headers, or both, from scoring.
ocarm* โ OCArm, an OCR pipeline built for historical Armenian printed periodicals (see the dataset card).
DeepSeek-OCR, Qwen-VL and Chandra were also evaluated but scored 50%+ CER and are not listed.
โ๏ธ How to submit results
Evaluate your model with the kit in the dataset's evaluation_kit/ folder.
Open this Space's Files tab, click results.csv โ Edit.
Add one row per model type: structure (aware / independent), model, then the four CER values in percent.
Click Open as a pull request and include a link to your model and your evaluation output.
The maintainers review the numbers and merge; the leaderboard updates automatically.