# Stage-1 document-extraction deps (HOST-side, OPTIONAL).
#
# These power scripts/extract-docs.py + extract-html.py (run on the host by
# scripts/extract-raw.sh / install-extract-cron.sh). Every backend is imported
# per-format and degrades gracefully — a missing dep just skips that format — so
# install only what you need. See docs/ingest-extraction.md.
#
#   python-docx  -> .docx        python-pptx -> .pptx
#   openpyxl     -> .xlsx        striprtf    -> .rtf
#   trafilatura  -> better HTML (extract-html.py falls back to the stdlib without it)
#
# System tools (not pip): poppler-utils (.pdf, `pdftotext`) and antiword/catdoc
# (legacy .doc) — `apt-get install poppler-utils antiword`.
python-docx>=1.1
python-pptx>=0.6.23
openpyxl>=3.1
striprtf>=0.0.26
trafilatura>=1.8     # optional; better HTML article extraction
