Skip to main content

Supported File Types and Document Limits

Supported file types, document volume limits, OCR, and semantic search scope in Unplex.

Written by Sabina Sokolowski

1. Supported File Types

Word documents (.docx, .doc): natively supported via the Word Add-in and direct upload — the primary format for document review workflows.

PDF (.pdf): fully supported. Text is extracted automatically; scanned PDFs are processed via the integrated OCR engine.

Excel (.xlsx, .xls): supported for structured data analysis and document review.

Email files (.eml, .msg): supported for review and analysis of correspondence.

Scanned documents and images: supported via integrated OCR; text is extracted, structured, and prepared for AI analysis automatically, with no manual pre-classification required.

For how each file type is processed on upload — including the specialized engine used for financial statement tables — see Uploading and Processing Documents.

2. Document Structure Handling

Document structure is detected automatically using AI, so semantic vectorization is calculated correctly without manual pre-classification. The system combines semantic vector embeddings, classical indexing, and LLM-based re-ranking, producing traceable results even for atypically formatted documents.

3. Document Volume Limits

Two distinct limits apply, governing different operations:

Per-file size: 50 MB per file is the governing constraint on individual document size (see Uploading and Processing Documents for full detail, including OCR page limits). Page count alone does not determine rejection; a large, image-heavy file reaches the 50 MB ceiling before an arbitrary page count would apply.

Per-analysis volume: up to 100 documents per Matrix analysis (standard tier). Higher volumes are supported via Enterprise configuration — contact support@unplex.ai.

4. Semantic Search Scope

Semantic search covers all documents actively stored in the workspace. Documents in connected cloud storage (SharePoint, OneDrive, Google Drive) are indexed once connected. The search is meaning-based rather than keyword-based, and operates across style, structure, and content.

Note: semantic search covers only documents explicitly added to or connected within the workspace. The full cloud storage archive is not automatically indexed.

5. In Development

PDFTools (advanced PDF manipulation and conversion) and a standalone OCR Engine are in development. DeepSign (digital signature workflow) is planned.

Did this answer your question?