Skip to main content

Supported File Types and Formats

The complete list of file formats Docupath accepts for upload, including document, image, spreadsheet, and structured data types

Docupath accepts a wide variety of document formats including PDFs, Microsoft Office documents, spreadsheets, presentations, images, text files, and structured data formats.

Each format has specific behaviour characteristics, and some formats deliver higher extraction accuracy than others. Understanding supported formats and their limitations is essential for optimizing document workflows and ensuring consistent data extraction quality.

How It Works

Complete List of Supported Formats

Docupath processes the following file types: Document Formats:

  • PDF (.pdf) - Preferred format; highest extraction accuracy

  • DOCX (.docx) - Microsoft Word; maintains text formatting and structure

  • DOC (.doc) - Legacy Microsoft Word; limited formatting preservation

  • HTML (.html) - Web documents; text extraction with minimal styling

  • TXT (.txt) - Plain text; minimal metadata extraction Presentation Formats:

  • PPTX (.pptx) - Microsoft PowerPoint; text extracted from slides

  • PPT (.ppt) - Legacy Microsoft PowerPoint; text extracted from slides Spreadsheet Formats:

  • XLSX (.xlsx) - Microsoft Excel; cell data and formulas extracted as line items

  • XLS (.xls) - Legacy Microsoft Excel; basic data extraction Image Formats:

  • JPG / JPEG (.jpg, .jpeg) - Standard image format; requires OCR processing

  • PNG (.png) - Portable network graphics; requires OCR processing

  • TIFF / TIF (.tiff, .tif) - Tagged image format; high-quality scans; requires OCR

  • HEIC (.heic) - Apple image format; automatically converted to JPG for processing Structured Data Formats:

  • XML (.xml) - Structured data; parsed as semi-structured content

  • JSON (.json) - Structured data; parsed as semi-structured content EDI Formats:

  • X12 - ANSI X12 electronic data interchange; parsed as structured content

  • EDL - Parsed as structured content All formats must meet the following constraints: maximum file size of 50MB, maximum 150 pages per document, and maximum 500 line items per document.

Category

Formats

Document

PDF, DOCX, DOC, HTML, TXT

Presentation

PPTX, PPT

Spreadsheet

XLSX, XLS

Image

JPG/JPEG, PNG, TIFF/TIF, HEIC

Structured Data

XML, JSON

EDI

X12, EDL

Format-Specific Behaviour

PDF Documents PDF is the recommended format for Docupath processing. PDFs with embedded text (searchable PDFs) achieve the highest extraction accuracy and fastest processing times. PDFs created from scanned images (image-based PDFs) require OCR processing and take longer to process (5-10 minutes vs. 1-3 minutes for text-based PDFs). Password-protected PDFs must be unlocked before upload; Docupath does not process encrypted PDFs. Fillable PDF forms are processed as standard documents; form fields are extracted as document fields.

Microsoft Word (DOCX/DOC) DOCX and DOC files are processed by extracting text content while preserving paragraph structure and heading hierarchy. Tables within Word documents are extracted as line item collections. Images embedded in Word documents are ignored during extraction; only text content is processed. Tracked changes and comments are not processed. DOCX files (newer format) achieve better accuracy than DOC files (legacy format) due to improved text extraction.

HTML Documents HTML files are processed by extracting all visible text content. HTML structure (headings, lists, tables) is preserved during extraction. Embedded JavaScript, CSS styling, and metadata tags are ignored. HTML files downloaded from websites may include navigation elements and footers; these are processed as regular content (not filtered out).

Plain Text (TXT) Plain text files are processed by extracting all content as a single text field. No structure is inferred; no paragraph or section detection occurs. Metadata such as timestamps or formatting is not extracted. Plain text files achieve perfect accuracy as there is no extraction ambiguity.

Microsoft PowerPoint (PPTX/PPT) PPTX and PPT files are processed by extracting text content from slides. Tables within slides are extracted as line item collections. Images, charts, and animations are ignored during extraction; only text content is processed. PPTX files (newer format) achieve better accuracy than PPT files (legacy format) due to improved text extraction.

Microsoft Excel (XLSX/XLS) Excel spreadsheets are processed by extracting data from the active worksheet. All cells in the active sheet are treated as potential data. Headers are extracted as field names, and data rows are extracted as line items (up to 500 per document). Formulas are calculated and values are extracted; formulas themselves are not captured. Multiple worksheets are not processed; only the active/first worksheet is used. Data exceeding 500 rows is truncated.

Image Formats (JPG, PNG, TIFF, HEIC) Image files undergo Optical Character Recognition (OCR) to extract text content. Image quality significantly impacts OCR accuracy; high-resolution scans (300+ DPI) achieve better accuracy than low-resolution images. Images must contain sufficient contrast between text and background. Handwritten content is supported but achieves lower accuracy than printed text. HEIC files are automatically converted to JPG format before processing. Color images, grayscale images, and black-and-white images are all processed; color is not required.

XML Data XML files are parsed and converted to structured data based on element and attribute names. The XML schema is not validated; malformed XML may produce unpredictable extraction results. Namespaces are preserved in element names. XML processing is useful for documents already in semi-structured format.

JSON Data JSON files are parsed and converted to structured data based on key and object names. The JSON schema is not validated; malformed JSON may produce unpredictable extraction results. Nested objects and arrays are preserved in the extracted structure. JSON processing is useful for documents already in semi-structured format.

EDI Formats (X12, EDL) X12 and EDL files are parsed as electronic data interchange (EDI) content and converted to structured data. EDI processing is useful for transactional documents already in a standardized machine format.

Format Recommendations and Best Practices

Optimal Formats by Document Type:

  • Invoices: PDF (searchable) - fastest processing, highest accuracy

  • Purchase Orders: PDF or XLSX - structured data extraction works well

  • Contracts: PDF (searchable) - accurate text extraction for clause identification

  • Bank Statements: PDF or XLSX - tabular data extracts cleanly

  • Scanned Documents: TIFF (high DPI) or PDF (image-based) - supported Quality Optimization:

  • Use searchable PDFs rather than image-based PDFs when possible. If converting documents to PDF, use a tool that preserves embedded text (e.g., office software export to PDF) rather than scanning.

  • For scanned documents, ensure 300+ DPI resolution. Lower resolution (96 DPI) may result in OCR errors.

  • Ensure high contrast between text and background. Faded, low-contrast documents produce lower OCR accuracy.

  • For documents with multiple languages, Docupath supports single-language documents best. Mixed-language documents may experience reduced accuracy.

Unsupported Formats

The following formats are not supported and are rejected at upload:

  • Video files (.mp4, .mov, .avi)

  • Audio files (.mp3, .wav)

  • Compressed files (.zip, .rar, .7z)

  • Executable files (.exe, .app)

  • Database files (.db, .sqlite) Attempting to upload unsupported formats results in an upload rejection with a clear message indicating which format is not supported.

Supported Configurations and Options

Format

Supported

Max File Size

OCR Required

Recommended Use

DOCX

Yes

50 MB

No

Word documents, reports

PPTX/PPT

Yes

50 MB

No

Presentations

XLSX

Yes

50 MB

No

Spreadsheet data

JPG/JPEG

Yes

50 MB

Yes

Scanned documents, photos

TIFF/TIF

Yes

50 MB

Yes

High-resolution scans

HTML

Yes

50 MB

No

Web documents

XML

Yes

50 MB

No

Structured data

JSON

Yes

50 MB

No

Structured data

X12

Yes

50 MB

No

EDI / structured data

EDL

Yes

50 MB

No

EDI / structured data

Other Technical Specifications

Specification

Value

Notes

Maximum Page Count

150 pages

All formats; calculated after conversion

PDF Text Extraction Speed

1-3 minutes

Searchable PDFs (embedded text)

Image OCR Processing

5-15 minutes

Depends on resolution and image quality

Excel Row Limit

500 rows

Excess rows are truncated

HTML DOM Depth

Unlimited

Entire DOM tree is processed

OCR Language Support

100+ languages

English, Spanish, French, German, Chinese, Japanese, Korean, Arabic, and more

OCR Accuracy (Low-Quality)

70-85%

Low-resolution or low-contrast images

Notes

  • Password-Protected PDFs: PDFs with encryption or password protection cannot be processed. Users must remove password protection before uploading.

  • Corrupted PDF Files: PDFs with corrupted metadata or internal structure may fail processing. Recommend re-exporting from the source application.

  • Scanned PDF Quality Variation: PDFs created by scanning physical documents have variable quality depending on scanner settings. Low-quality scans (96 DPI, poor lighting) produce lower extraction accuracy.

  • Multi-Column Documents: Documents with multi-column layouts (common in PDFs and scans) may have text extracted in reading order errors. Columns may be concatenated or reordered.

  • Handwritten Content: Handwritten text in documents or forms is partially supported but achieves 70-80% accuracy depending on handwriting legibility. Printed text is always more accurate.

  • Complex Tables: Tables with merged cells, nested tables, or irregular structure may extract incorrectly. Simple, well-structured tables extract most accurately.

  • Document Rotation: Images or PDFs with pages rotated 90/180 degrees may extract incorrectly. Ensure pages are oriented normally before upload.

  • Mixed Languages: Documents containing multiple languages may experience reduced accuracy. Single-language documents are recommended for best results.

  • Excel Formulas: Excel formulas are evaluated but the formula text is not captured. Only the calculated result is extracted.

  • Legacy Document Formats: DOC and XLS files (pre-2007 Microsoft formats) achieve lower accuracy than DOCX and XLSX. Conversion to modern formats is recommended.

Did this answer your question?