CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
-
Updated
Dec 31, 2019
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
This repository contains a 403 images dataset for table detection in documents.
Creates synthetic degraded image documents that could be used to train Neural Networks
Total Text Dataset. It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind.
Distorted Document Images dataset (DDI-100).
Scripts and results from our OCR roundup, available on Source
Official implementation of SynthTIGER (Synthetic Text Image GEneratoR) ICDAR 2021
EATEN: Entity-aware Attention for Single Shot Visual Text Extraction
A synthetic data generator for text recognition
A repository with anonymized invoices
Generate text images for training deep learning ocr model
EATEN: Entity-aware Attention for Single Shot Visual Text Extraction
ScrabbleGAN: Semi-Supervised Varying Length Handwritten Text Generation (CVPR20)
Ground truth line annotations for the Berliner Börsen-Zeitung
This is a simple project to generate simple cropped images with characters. You can generate with Chinese or English characters. Backgrounds are also allowed. Medical bills simulation are also included.
Master's thesis work as a part of M2(Advanced Robotics) @ Centrale Nantes
Parliamentary Bills Classification using Document Level Embedding and Bidirectional LongShort-Term Memory
To associate your repository with the aniketdataset topic, visit your repo's landing page and select "manage topics."