CRISP
LUDDY HACKATHON / CASE 2 / 2026
CRISP.
Printed pages in. Compressed bytes out.
COMPRESSED RECOGNITION FOR IMAGE STORAGE PIPELINE
TEAM / SURYA · ANISH · LAHARI
STAGE 01
OCR Pipeline
IMAGE > TEXT
01
Denoise
CNN
02
Segment
CV
03
Recognize
CNN
90.6%
CHARACTER ACCURACY
111ms
MEAN STAGE 1 LATENCY
CAVEAT — DOMAIN GAP
Trained and benchmarked on EMNIST. Real-world smartphone scans differ in lighting, skew, font, and resolution — so live accuracy is lower than the benchmark.
STAGE 02
Adaptive Huffman
VITTER'S ALGORITHM V
2A
TEXT > BYTES
Compress.
Single pass, dynamic Huffman. Tree built as characters stream in. No dictionary transmitted.
1.88x
MEAN COMPRESSION RATIO
2B
BYTES > TEXT
Decompress.
Mirrors the encoder bit by bit. Rebuilds the same tree as it reads. 100% lossless round trip.
20/ 20
ROUND TRIPS VERIFIED
REAL-WORLD & VIABILITY
Why CRISP Ships
UNDER 200 MS / PAGE
01 / THE NEED
Scan-heavy archives.
Legal, medical, historical. Millions of printed pages stored as raw images.
Under 200 ms / page
02 / BUILT TO DEPLOY
Two containers, not a notebook.
Dockerized FastAPI services. Tested. Warmed. GPU preferred, CPU fallback.
20 / 20 lossless round trips
03 / ROOM TO GROW
Modular upgrade path.
U-Net denoiser. True-case recognizer. Arithmetic coder. Wire contract unchanged.
3 independent upgrade axes
CRISP / 2026
Q & A
Thank you.
Questions welcome.
SURYA ANISH LAHARI