Add local receipt OCR + grouping (OCR/)

Two local, GPU-backed tools that talk to a local ollama server (no cloud,
no API key):

- receipt_ocr.py: receipt image -> ordered list of rows of text, via
  Qwen2.5-VL 7B. Zero deps (stdlib + ollama HTTP API).
- receipt_group.py: OCR rows -> structured receipt (store, line items with
  multi-row details merged, total, metadata), via Mistral Small 24B. Detail
  rows are attached by arithmetic reconciliation, not position; a prompt
  completeness clause guards against dropped rows; a deterministic --json
  reconciliation pass audits item math, subtotal/total, and unit count.

README documents setup, an ollama operating runbook, and design rationale.
Includes the ALDI example image + its OCR rows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Jean-Michel Tremblay 2026-06-22 22:47:09 -04:00
parent e94a17160b
commit 344663d87d
5 changed files with 725 additions and 0 deletions

263
OCR/README.md Normal file
View file

@ -0,0 +1,263 @@
I want to create an OCR feature (ollama, whatever) that's
1) local (assume I have a GPU with 24GB VRAM)
2) works on receipts (scanned or digital)
3) given a receipt, return an ordered list of rows of text in that receipt
4) example /home/jm/programming/HSAmanager/OCR/example/data/aldi.jpg
---
## Getting started (working prototype)
Local, GPU-backed, no cloud / no API key. Engine: **Qwen2.5-VL 7B** served by
**ollama**. The model transcribes the receipt top-to-bottom and returns one
string per physical row, columns joined left-to-right.
### One-time setup (already done on this machine)
- ollama binary installed (no root) at `~/.local/bin/ollama`.
- Model pulled: `ollama pull qwen2.5vl:7b` (~6 GB).
### Run
```bash
# 1. start the server (leave running; ~6 GB VRAM when the model loads)
~/.local/bin/ollama serve &
# 2. transcribe a receipt -> ordered, numbered rows
python3 OCR/receipt_ocr.py OCR/example/data/aldi.jpg
# raw JSON {"rows": [...]} instead of numbered text
python3 OCR/receipt_ocr.py OCR/example/data/aldi.jpg --json
```
No third-party Python deps — `receipt_ocr.py` uses the stdlib and the ollama
HTTP API. Override with `OLLAMA_MODEL` / `OLLAMA_HOST` env vars.
### Notes / next steps
- Works on the ALDI example with near-perfect line ordering and column joining.
- Caveat: the model sometimes *expands* thermal-print abbreviations
(`Fire Roasted Tom` → `Fire Roasted Tomatoes`). Helpful for readability but
not byte-for-byte faithful; tighten the prompt if you need verbatim text.
- Accuracy upgrade: `qwen2.5vl:32b` (q4 ≈ 20 GB, fits the 3090 but tighter/slower).
- This is local-only and complementary to `../receiptscan/`, which uses the
Claude cloud API for line-item categorization + annotated rendering.
---
## Operating ollama (runbook)
ollama is the local model server both scripts talk to over HTTP
(`http://127.0.0.1:11434`). It must be running before you call either script.
Installed **without root** at `~/.local/bin/ollama` (the binary was extracted
from the `ollama-linux-amd64.tar.zst` GitHub release; `~/.local/bin` is on PATH,
so plain `ollama` works too).
### Start / check / stop
```bash
# Start the server (does NOT survive a reboot or terminal close — restart it):
ollama serve & # or: nohup ollama serve > /tmp/ollama.log 2>&1 &
# Is it up? (prints a JSON version string if so)
curl -fsS http://127.0.0.1:11434/api/version
# Is the process alive?
pgrep -af "ollama serve"
# Stop it:
pkill -f "ollama serve"
```
After a reboot the server is **not** running — just `ollama serve &` again. The
pulled models persist on disk (under `~/.ollama`), so you don't re-download them.
### Models
```bash
ollama list # models on disk
ollama ps # models currently loaded in VRAM (and how long until idle-unload)
ollama pull <model> # download a model, e.g. ollama pull qwen2.5:14b
ollama rm <model> # delete a model to reclaim disk
```
Currently pulled: `qwen2.5vl:7b` (OCR), `mistral-small3.2:24b` (grouping, default),
plus `qwen2.5:7b` / `qwen2.5:14b` (earlier grouping experiments — `ollama rm`
them if you want the disk back).
### VRAM / GPU notes
- A model loads into VRAM on first request and **auto-unloads after ~5 min idle**
(so `ollama ps` is often empty — that's normal; the next call reloads it in a
few seconds).
- Only run one large model at a time on the 24 GB 3090. The OCR model (~6 GB) and
the grouping model (~15 GB) can coexist, but bumping grouping to
`qwen2.5:32b`/`qwen2.5vl:32b` (~20 GB) leaves little room — expect a reload
when switching between OCR and grouping.
- Force-unload now (free VRAM without stopping the server):
`ollama stop <model>`.
- Keep a model resident longer / shorter via the request `keep_alive` field, or
globally with `OLLAMA_KEEP_ALIVE` (e.g. `OLLAMA_KEEP_ALIVE=30m ollama serve`).
### Pointing the scripts elsewhere
Both scripts honor `OLLAMA_HOST` and `OLLAMA_MODEL`:
```bash
OLLAMA_MODEL=qwen2.5:14b python3 OCR/receipt_group.py # try another model
OLLAMA_HOST=http://other-box:11434 python3 OCR/receipt_ocr.py img.jpg # remote server
```
### Logs / troubleshooting
- If a script hangs or errors connecting, the server is probably down — check
with the `curl` above and restart.
- Startup logs (GPU discovery, VRAM, errors) go wherever you redirected `serve`
(e.g. `/tmp/ollama.log`); tail that if a model fails to load.
- "model not found" → you haven't `ollama pull`ed it (or typo'd the tag); see
`ollama list`.
# grouping the OCR results
next task is ... given an OCR output, like:
```text
1 ALDI
2 Store #145
3 1501 Rockville Pike
4 Rockville, MD
5 https://help.aldi.us
6 Your cashier today was Wendy
7 382175 Unsalted Peanuts 2.29 FA
8 382175 Unsalted Peanuts 2.29 FA
9 384773 4 lb. Sugar 2.79 FA
10 384773 4 lb. Sugar 2.79 FA
11 382437 Fire Roasted Tomatoes 2.30 FA
12 2 @ 1.15
13 371607 Canned Cat Food 2.28 NB
14 4 @ 0.57
15 371607 Canned Cat Food 0.57 NB
16 416787 1% Milk, Gallon 3.38 FA
17 343557 Protein Powder 18.49 FA
18 416645 Chocolate Milk 1.91 FA
19 382260 Plain NF Greek Yogurt 2.79 FA
20 382260 Plain NF Greek Yogurt 2.79 FA
21 297956 PureandSimple Bars 3.99 FA
22 356508 Broccoli Crowns 2.21 FA
23 1.17 lb x 1.89/lb
24 469529 Corn Tortillas 1.95 FA
25 609748 Assorted Cashews 6.49 FA
26 634568 Chicken Skewers 7.49 FA
27 382473 Shredded Mozzarella 3.29 FA
28 341878 Yellow Onions 1.85 FA
29 535290 ABF B/S Thighs 6.61 FA
30 272135 Salmon Portions 9.35 FA
31 272135 Salmon Portions 8.54 FA
32 382310 Feta Crumbles 1.29 FA
33 356508 Broccoli Crowns 2.53 FA
34 1.34 lb x 1.89/lb
35 262137 Stuffed Olives 2.89 FA
36 356490 Bagged Avocados 2.99 FA
37 356607 Red Delic Apples 2.75 FA
38 356527 Celery 1.89 FA
39 382653 Indian Sauces 3.69 FA
40 356628 Seedless Cucumber 0.89 FA
41 356684 WildTwist Apf LRW 5.44 FA
42 (G) 3.10 lb - (T) 0.06 lb
43 (N) 3.04 lb x 1.79/lb
44 282119 Pineapples 1.89 FA
45 356419 Mangoes 3.80 FA
46 4 @ 0.95
47 356691 Zucchini 2.01 FA
48 1.69 lb x 1.19/lb
49 356522 Cantaloupe 1.89 FA
50 356504 Blueberries 9.95 FA
51 5 @ 1.99
52 388137 Large Eggs 2.92 FA
53 2 @ 1.46
54 262747 Bananas LRW 1.45 FA
55 (G) 2.97 lb - (T) 0.01 lb
56 (N) 2.96 lb x 0.49/lb
57 262747 Bananas LRW 1.13 FA
58 (G) 2.32 lb - (T) 0.01 lb
59 (N) 2.31 lb x 0.49/lb
60 356427 Multi-Peppers 3pk. 2.69 FA
61 341876 Red Grapes LRW 5.48 FA
62 (G) 3.96 lb - (T) 0.02 lb
63 (N) 3.94 lb x 1.39/lb
64 479744 Campari Tomatoes 2.99 FA
65 479744 Campari Tomatoes 2.99 FA
66 569269 Protein Bread 3.99 FA
67 343989 Family AsstCookie 1.99 FA
68 356646 Strawberries 2.09 FA
69 356646 Strawberries 2.09 FA
70 356646 Strawberries 2.09 FA
71 VISA 172.42
72 ************0294 ONLINE
73 06/21/26 12:43 Ref/Seq # 0
74 Auth# 00401D
75 AID A000000031010
76 TVR 0000000000
77 IAD 06021203A00000
78 TSI 0000 ARC 00 EntryMode 07
79 ++APPROVED++
80 SUBTOTAL 172.24
81 B-Taxable @6.00% 0.18
82 A-Taxable @0.00% 0.00
83 AMOUNT DUE 172.42
84 TOTAL $172.42
85 60 ITEMS
86 Credit Card $ 172.42
87 *7367 L411/005/064 06/21/26 12:43PM
88 ************
89 Sign up for ALDI emails
90 for a sneak peek on the weekly ad!
91 www.aldi.us/signup
```
write a LLM call, still using my local ollama (you can switch the model). I want the output to be an object roughly like that:
store name:
store branch:
store address:
list of rows that belong together as one item purchased
total in dollars.
(anything else is relevant here)?
you can hardcode the lines for this test and let's call the LLM to see what it does
the output needs to be a new script (receipt_group.py) that (for now) read this examle as is from an example file that you will create, and with the correct prompt organizes the information correctly.
## Grouping — implemented
`receipt_group.py` reads OCR rows (default
`example/data/aldi_ocr.txt`) and asks a local ollama text model to reconstruct
the receipt: store name/branch/address, line items (merging multi-row items
such as weighed produce and `qty @ price` lines), total, and metadata
(cashier, datetime, payment, subtotal, tax, item count).
```bash
python3 OCR/receipt_group.py # groups only: each item + its rows
python3 OCR/receipt_group.py OCR/example/data/aldi_ocr.txt --json # full object + reconciliation
```
- Default model: `mistral-small3.2:24b` (~15 GB). Override with `OLLAMA_MODEL`.
Tried qwen2.5:7b (many merge errors) and qwen2.5:14b (~90%); Mistral Small got
the grouping right. Swapping models is just the env var — no code change.
- **Why an LLM groups, not a regex**: detecting item vs detail rows is trivial,
but deciding *which* item a detail row attaches to is not — wrapped
descriptions, discount lines, and qty/weight lines can sit above or below the
item, so a positional rule breaks. The prompt tells the model to attach each
detail row to the item its **math** reconciles with (count × unit, weight ×
rate = the item's charged price), never by a fixed above/below position.
- **Completeness clause** in the prompt: every input row must land in exactly one
item's `rows`, copied verbatim, and the model must walk the input top-to-bottom
to confirm nothing was dropped. This fixed the model silently swallowing Large
Eggs' `2 @ 1.46` line. Prompting lowers the odds of a drop but cannot guarantee
it — hence the backstop below.
- **Deterministic reconciliation pass** (the part you *can* trust), shown in
`--json` under `validation`: Python recomputes each item from its detail math,
sums items vs the printed subtotal, checks total = subtotal + tax, and sums
unit quantities vs the printed `N ITEMS`. When the model slips (e.g. dropping a
`qty @` line), the unit-count check flags it (59 vs 60) even though dollars
still balance.

BIN
OCR/example/data/aldi.jpg Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 438 KiB

View file

@ -0,0 +1,91 @@
1 ALDI
2 Store #145
3 1501 Rockville Pike
4 Rockville, MD
5 https://help.aldi.us
6 Your cashier today was Wendy
7 382175 Unsalted Peanuts 2.29 FA
8 382175 Unsalted Peanuts 2.29 FA
9 384773 4 lb. Sugar 2.79 FA
10 384773 4 lb. Sugar 2.79 FA
11 382437 Fire Roasted Tomatoes 2.30 FA
12 2 @ 1.15
13 371607 Canned Cat Food 2.28 NB
14 4 @ 0.57
15 371607 Canned Cat Food 0.57 NB
16 416787 1% Milk, Gallon 3.38 FA
17 343557 Protein Powder 18.49 FA
18 416645 Chocolate Milk 1.91 FA
19 382260 Plain NF Greek Yogurt 2.79 FA
20 382260 Plain NF Greek Yogurt 2.79 FA
21 297956 PureandSimple Bars 3.99 FA
22 356508 Broccoli Crowns 2.21 FA
23 1.17 lb x 1.89/lb
24 469529 Corn Tortillas 1.95 FA
25 609748 Assorted Cashews 6.49 FA
26 634568 Chicken Skewers 7.49 FA
27 382473 Shredded Mozzarella 3.29 FA
28 341878 Yellow Onions 1.85 FA
29 535290 ABF B/S Thighs 6.61 FA
30 272135 Salmon Portions 9.35 FA
31 272135 Salmon Portions 8.54 FA
32 382310 Feta Crumbles 1.29 FA
33 356508 Broccoli Crowns 2.53 FA
34 1.34 lb x 1.89/lb
35 262137 Stuffed Olives 2.89 FA
36 356490 Bagged Avocados 2.99 FA
37 356607 Red Delic Apples 2.75 FA
38 356527 Celery 1.89 FA
39 382653 Indian Sauces 3.69 FA
40 356628 Seedless Cucumber 0.89 FA
41 356684 WildTwist Apf LRW 5.44 FA
42 (G) 3.10 lb - (T) 0.06 lb
43 (N) 3.04 lb x 1.79/lb
44 282119 Pineapples 1.89 FA
45 356419 Mangoes 3.80 FA
46 4 @ 0.95
47 356691 Zucchini 2.01 FA
48 1.69 lb x 1.19/lb
49 356522 Cantaloupe 1.89 FA
50 356504 Blueberries 9.95 FA
51 5 @ 1.99
52 388137 Large Eggs 2.92 FA
53 2 @ 1.46
54 262747 Bananas LRW 1.45 FA
55 (G) 2.97 lb - (T) 0.01 lb
56 (N) 2.96 lb x 0.49/lb
57 262747 Bananas LRW 1.13 FA
58 (G) 2.32 lb - (T) 0.01 lb
59 (N) 2.31 lb x 0.49/lb
60 356427 Multi-Peppers 3pk. 2.69 FA
61 341876 Red Grapes LRW 5.48 FA
62 (G) 3.96 lb - (T) 0.02 lb
63 (N) 3.94 lb x 1.39/lb
64 479744 Campari Tomatoes 2.99 FA
65 479744 Campari Tomatoes 2.99 FA
66 569269 Protein Bread 3.99 FA
67 343989 Family AsstCookie 1.99 FA
68 356646 Strawberries 2.09 FA
69 356646 Strawberries 2.09 FA
70 356646 Strawberries 2.09 FA
71 VISA 172.42
72 ************0294 ONLINE
73 06/21/26 12:43 Ref/Seq # 0
74 Auth# 00401D
75 AID A000000031010
76 TVR 0000000000
77 IAD 06021203A00000
78 TSI 0000 ARC 00 EntryMode 07
79 ++APPROVED++
80 SUBTOTAL 172.24
81 B-Taxable @6.00% 0.18
82 A-Taxable @0.00% 0.00
83 AMOUNT DUE 172.42
84 TOTAL $172.42
85 60 ITEMS
86 Credit Card $ 172.42
87 *7367 L411/005/064 06/21/26 12:43PM
88 ************
89 Sign up for ALDI emails
90 for a sneak peek on the weekly ad!
91 www.aldi.us/signup

284
OCR/receipt_group.py Executable file
View file

@ -0,0 +1,284 @@
#!/usr/bin/env python3
"""Group raw receipt-OCR rows into a structured object with a local LLM.
Takes the text output of receipt_ocr.py (one row per line, optionally prefixed
with a line number) and asks a local ollama model to reconstruct the receipt:
store identity, the line items (merging multi-row items such as weighed produce
and "qty @ price" lines), the total, and any other useful metadata.
Local-only: no cloud, no API key. Uses a text model (default qwen2.5:7b).
Usage:
python3 receipt_group.py [ocr_rows.txt] # pretty summary
python3 receipt_group.py ocr_rows.txt --json # raw structured JSON
Env:
OLLAMA_HOST default http://127.0.0.1:11434
OLLAMA_MODEL default qwen2.5:7b
"""
import json
import os
import re
import sys
import urllib.request
DEFAULT_INPUT = os.path.join(os.path.dirname(__file__), "example", "data", "aldi_ocr.txt")
HOST = os.environ.get("OLLAMA_HOST", "http://127.0.0.1:11434").rstrip("/")
MODEL = os.environ.get("OLLAMA_MODEL", "mistral-small3.2:24b")
PROMPT = """You are given the OCR transcription of a single retail receipt, one \
physical row per line. Each line may start with a row number (e.g. "23 ...") \
ignore the row number itself, it is not part of the receipt text.
Reconstruct the receipt as structured data. Rules:
- A purchased ITEM begins on a row that has a product code (a run of digits) and \
a price. EVERY row that begins with a product code starts a NEW item never \
attach a code-bearing row to the item above it, even when the name and price are \
identical to the row above. Two identical consecutive code rows are TWO separate \
purchases, each its own entry. Do not merge or de-duplicate them.
- A DETAIL row is any row that does NOT begin with a product code, e.g.:
* quantity lines like "2 @ 1.15" (count @ unit price)
* weight lines like "1.17 lb x 1.89/lb"
* gross/tare/net lines like "(G) 3.10 lb - (T) 0.06 lb" and "(N) 3.04 lb x 1.79/lb"
Attach each detail row to the item it actually describes. Decide that by the \
MATH, not by position: count x unit price, or weight x rate, equals that item's \
charged price. (A gross/tare line that has no price of its own stays with the \
net line it accompanies.) The describing item is usually adjacent, but verify by \
reconciliation rather than assuming it is the one above or below.
- The item's "price" is the dollar amount printed on the item's own (first) row \
the amount actually charged. Keep two decimals.
- Put the exact original row strings that belong to each item in its "rows" array, \
in order, WITHOUT the leading row number.
- Header rows (store name, branch, address, website, cashier) are NOT items.
- Payment, auth, subtotal, tax, total, item-count and marketing/footer rows are \
NOT items pull the useful values into the metadata fields instead.
- Do not invent values. If something is not present, omit it.
COMPLETENESS this is critical:
- Every item row AND every detail row from the input must appear in exactly ONE \
item's "rows" array, copied VERBATIM (same digits, same spacing, same price). \
You may not drop, skip, summarize, or rewrite any row.
- NEVER discard a quantity line ("2 @ 1.46"), a weight line ("1.17 lb x 1.89/lb"), \
or a gross/tare/net line. Attach each to the item its math reconciles with.
- Before you answer, walk the input top to bottom and confirm that every row is \
either placed in an item's "rows" or is a header/payment/footer row that you \
deliberately excluded. No item or detail row may be missing.
Return STRICT JSON only, matching the requested schema."""
SCHEMA = {
"type": "object",
"properties": {
"store_name": {"type": "string"},
"store_branch": {"type": "string"},
"store_address": {"type": "string"},
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"code": {"type": "string"},
"price": {"type": "number"},
"rows": {"type": "array", "items": {"type": "string"}},
},
"required": ["name", "price", "rows"],
},
},
"total": {"type": "number"},
"metadata": {
"type": "object",
"properties": {
"cashier": {"type": "string"},
"purchase_datetime": {"type": "string"},
"payment_method": {"type": "string"},
"subtotal": {"type": "number"},
"tax": {"type": "number"},
"item_count": {"type": "integer"},
},
},
},
"required": ["store_name", "items", "total"],
}
def group(rows_text):
payload = {
"model": MODEL,
"messages": [
{"role": "user", "content": PROMPT + "\n\n--- RECEIPT OCR ---\n" + rows_text},
],
"stream": False,
"format": SCHEMA,
"options": {"temperature": 0, "num_ctx": 8192},
}
req = urllib.request.Request(
f"{HOST}/api/chat",
data=json.dumps(payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=600) as resp:
body = json.load(resp)
return json.loads(body["message"]["content"])
# --- Deterministic reconciliation -------------------------------------------
# The LLM decides the grouping; arithmetic decides whether to trust it. Detail
# rows carry their own math, so we recompute each item's price from its details
# and balance the whole receipt against the printed subtotal / total / count.
QTY_RE = re.compile(r"(\d+)\s*@\s*\$?(\d+(?:\.\d+)?)") # "4 @ 0.57"
WEIGHT_RE = re.compile(r"(\d+(?:\.\d+)?)\s*lb\s*[xX]\s*\$?(\d+(?:\.\d+)?)\s*/\s*lb") # "1.17 lb x 1.89/lb"
TOL = 0.02 # cents of slack for rounding on weighed items
def check_item(it):
"""Reconcile one item from its detail rows.
Returns (status, qty, note) where status is 'ok' | 'warn' | 'none'.
'none' = no math-bearing detail row, so there's nothing to verify.
"""
price = float(it.get("price", 0))
qty = 1
expected = None
note = ""
for row in it.get("rows", [])[1:]: # skip the anchor row; details only
m = QTY_RE.search(row)
if m:
qty = int(m.group(1))
expected = round(qty * float(m.group(2)), 2)
note = f"{m.group(1)} @ {m.group(2)} = {expected:.2f}"
continue
w = WEIGHT_RE.search(row)
if w:
expected = round(float(w.group(1)) * float(w.group(2)), 2)
note = f"{w.group(1)} lb x {w.group(2)}/lb = {expected:.2f}"
if expected is None:
return ("none", qty, "")
if abs(expected - price) <= TOL:
return ("ok", qty, note)
return ("warn", qty, f"{note} != charged {price:.2f}")
def validate(obj):
items = obj.get("items", [])
meta = obj.get("metadata", {}) or {}
per_item = [check_item(it) for it in items]
items_sum = round(sum(float(it.get("price", 0)) for it in items), 2)
unit_count = sum(q for _, q, _ in per_item)
subtotal = meta.get("subtotal")
tax = meta.get("tax") or 0
total_printed = obj.get("total")
count_printed = meta.get("item_count")
report = {
"per_item": [{"status": s, "qty": q, "note": n} for s, q, n in per_item],
"items_sum": items_sum,
"subtotal_printed": subtotal,
"subtotal_ok": subtotal is not None and abs(items_sum - subtotal) <= TOL,
"total_computed": round(items_sum + tax, 2),
"total_printed": total_printed,
"total_ok": total_printed is not None and abs(items_sum + tax - total_printed) <= TOL,
"unit_count": unit_count,
"unit_count_printed": count_printed,
"unit_count_ok": count_printed is not None and unit_count == count_printed,
}
return report
# --- Rendering ----------------------------------------------------------------
MARK = {"ok": "", "warn": "", "none": " "}
def yn(ok):
return "" if ok else ""
def pretty(obj, report):
out = []
out.append(f"Store: {obj.get('store_name', '')}")
if obj.get("store_branch"):
out.append(f"Branch: {obj['store_branch']}")
if obj.get("store_address"):
out.append(f"Address: {obj['store_address']}")
out.append("")
items = obj.get("items", [])
checks = report["per_item"]
out.append(f"Items ({len(items)}):")
for it, chk in zip(items, checks):
code = it.get("code", "")
head = f" {MARK[chk['status']]} {it.get('name', '')} ${it.get('price', 0):.2f}"
if code:
head += f" [{code}]"
out.append(head)
for extra in it.get("rows", [])[1:]: # detail rows only
out.append(f" · {extra}")
if chk["status"] == "warn":
out.append(f"{chk['note']}")
out.append("")
# Receipt-level balance
out.append("Reconciliation:")
out.append(
f" Items sum: ${report['items_sum']:.2f}"
+ (
f" (printed subtotal ${report['subtotal_printed']:.2f}) {yn(report['subtotal_ok'])}"
if report["subtotal_printed"] is not None
else ""
)
)
if report["total_printed"] is not None:
out.append(
f" + tax/total: ${report['total_computed']:.2f}"
f" (printed TOTAL ${report['total_printed']:.2f}) {yn(report['total_ok'])}"
)
if report["unit_count_printed"] is not None:
out.append(
f" Unit count: {report['unit_count']}"
f" (printed {report['unit_count_printed']} ITEMS) {yn(report['unit_count_ok'])}"
)
meta = obj.get("metadata", {})
if meta:
out.append("")
out.append("Metadata:")
for k, v in meta.items():
out.append(f" {k}: {v}")
return "\n".join(out)
def main():
args = [a for a in sys.argv[1:] if not a.startswith("--")]
as_json = "--json" in sys.argv
path = args[0] if args else DEFAULT_INPUT
if not os.path.exists(path):
sys.exit(f"input not found: {path}")
with open(path, encoding="utf-8") as f:
rows_text = f.read()
obj = group(rows_text)
if as_json:
obj["validation"] = validate(obj)
print(json.dumps(obj, ensure_ascii=False, indent=2))
else:
# Just the groups: each purchased item, with the rows that belong together.
items = obj.get("items", [])
width = len(str(len(items)))
for i, it in enumerate(items, 1):
rows = it.get("rows", [])
head = rows[0] if rows else f"{it.get('name', '')} {it.get('price', 0):.2f}"
print(f"{str(i).rjust(width)} {head}")
for extra in rows[1:]:
print(f"{' ' * width} {extra}")
if __name__ == "__main__":
main()

87
OCR/receipt_ocr.py Executable file
View file

@ -0,0 +1,87 @@
#!/usr/bin/env python3
"""Local receipt OCR: given a receipt image, print an ordered list of rows of text.
Uses a local Qwen2.5-VL model served by ollama. No cloud calls, no API key.
Zero third-party dependencies talks to the ollama HTTP API with the stdlib.
Usage:
python3 receipt_ocr.py [image.jpg] # pretty, numbered rows
python3 receipt_ocr.py image.jpg --json # raw JSON {"rows": [...]}
Env:
OLLAMA_HOST default http://127.0.0.1:11434
OLLAMA_MODEL default qwen2.5vl:7b
"""
import base64
import json
import os
import sys
import urllib.request
DEFAULT_IMAGE = os.path.join(os.path.dirname(__file__), "example", "data", "aldi.jpg")
HOST = os.environ.get("OLLAMA_HOST", "http://127.0.0.1:11434").rstrip("/")
MODEL = os.environ.get("OLLAMA_MODEL", "qwen2.5vl:7b")
PROMPT = (
"You are an OCR engine for retail receipts. Transcribe EVERY line of text on "
"this receipt, from top to bottom, exactly as printed. Each physical row of the "
"receipt becomes one string in the output, with its columns joined left-to-right "
"and separated by single spaces (e.g. an item code, its name, and its price all "
"go in the same row). Do not merge or split rows, do not reorder, do not add "
"commentary, do not correct spelling, do not invent text. Preserve numbers and "
"punctuation verbatim. Return strict JSON only."
)
# Structured-output schema: forces the model to return {"rows": ["...", ...]}.
SCHEMA = {
"type": "object",
"properties": {"rows": {"type": "array", "items": {"type": "string"}}},
"required": ["rows"],
}
def ocr(image_path):
with open(image_path, "rb") as f:
b64 = base64.b64encode(f.read()).decode("ascii")
payload = {
"model": MODEL,
"messages": [{"role": "user", "content": PROMPT, "images": [b64]}],
"stream": False,
"format": SCHEMA,
"options": {"temperature": 0, "num_ctx": 8192},
}
req = urllib.request.Request(
f"{HOST}/api/chat",
data=json.dumps(payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=600) as resp:
body = json.load(resp)
content = body["message"]["content"]
return json.loads(content)["rows"]
def main():
args = [a for a in sys.argv[1:] if not a.startswith("--")]
as_json = "--json" in sys.argv
image_path = args[0] if args else DEFAULT_IMAGE
if not os.path.exists(image_path):
sys.exit(f"image not found: {image_path}")
rows = ocr(image_path)
if as_json:
print(json.dumps({"rows": rows}, ensure_ascii=False, indent=2))
else:
width = len(str(len(rows)))
for i, row in enumerate(rows, 1):
print(f"{str(i).rjust(width)} {row}")
if __name__ == "__main__":
main()