Configuration

Configure Smole to fit your needs.

How the Pipeline Works

The Smole pipeline processes documents in two stages:

1. Conversion

Your document (PDF, DOCX, image, etc.) is converted to clean Markdown text. This step handles OCR for scanned documents, extracts tables, and preserves document structure. The output is a text representation that the AI can understand.

2. Extraction

Smole processes the Markdown content alongside your JSON schema, using the document context to extract structured data that matches your schema. Fields not found in the document are returned as null (see Schemas). The output is validated JSON matching your defined structure.

When you submit a pipeline job, both steps run automatically. You poll for status until the job completes, then retrieve your extracted JSON data. See Guides for a complete walkthrough.

Pipeline Parameters

When calling POST /api/pipeline/file (see API Reference), you can include these parameters as form fields:

Required Parameters

ParameterDescription
fileThe document file to process (multipart file upload)
schemaIdUUID of the schema to use for extraction. Create schemas via POST /api/schemas. To run a workflow, send workflowId instead: the workflow sets the preparation options, and sending them with it returns 400

Processing Options

ParameterDefaultDescription
processingModenormalOptional, accepted for compatibility. normal and fast are processed the same way at the same price.
maxPagesallMaximum number of pages to process. Useful for large documents where you only need the first few pages
preserveTablestruePreserve table structure as Markdown tables
preserveLinkstruePreserve hyperlinks in the converted Markdown
extractImagesfalseExtract embedded images and include descriptions
stripNavigationtrueRemove navigation elements (headers, footers, menus) from web pages
stripAdstrueRemove advertisement content from web pages

Supported Formats

Smole supports the following document formats:

.pdfPDF
.docxDOCX
.docDOC
.xlsxXLSX
.xlsXLS
.pptxPPTX
.pptPPT
.txtTXT
.htmlHTML
.rtfRTF
.odtODT
.csvCSV

Rate Limits

Default rate limits per user:

LimitValue
Requests per minute60
Free documents per day5

Starter includes 5,000 documents per month. Each document uses 1 document unit, whatever its page count.