Configuration
Configure Smole to fit your needs.
How the Pipeline Works
The Smole pipeline processes documents in two stages:
1. Conversion
Your document (PDF, DOCX, image, etc.) is converted to clean Markdown text. This step handles OCR for scanned documents, extracts tables, and preserves document structure. The output is a text representation that the AI can understand.
2. Extraction
Smole processes the Markdown content alongside your JSON schema, using the document context to extract structured data that matches your schema. Fields not found in the document are returned as null (see Schemas). The output is validated JSON matching your defined structure.
When you submit a pipeline job, both steps run automatically. You poll for status until the job completes, then retrieve your extracted JSON data. See Guides for a complete walkthrough.
Pipeline Parameters
When calling POST /api/pipeline/file (see API Reference), you can include these parameters as form fields:
Required Parameters
Processing Options
Supported Formats
Smole supports the following document formats:
Rate Limits
Default rate limits per user:
Starter includes 5,000 documents per month. Each document uses 1 document unit, whatever its page count.