Guides
Learn how to use Smole's features effectively.
Processing Documents
Upload documents with a schema ID to extract structured data. See Schemas for how to create and manage schemas, and Configuration for processing modes and supported document options.
curl -X POST https://api.smole.tech/api/pipeline/file \
-H "X-API-Key: ak_your_api_key" \
-F "file=@receipt.pdf" \
-F "schemaId=your_schema_id"The response includes a job ID. Poll /api/pipeline/:id until the status is completed. See API Reference for all pipeline endpoints.
Checking Job Status
Jobs go through several statuses during processing:
pending- Job queued for processingconverting- Document being converted to Markdownextracting- AI extracting data from Markdowncompleted- Extraction finished, result availablefailed- Error occurred, check error field
Understanding Results
When a job completes, the result field contains the extracted JSON matching your schema. Fields that couldn't be found in the document are returned as null. See Missing Data and Null Values for details.
{
"status": "completed",
"result": {
"companyName": "Acme Corp",
"invoiceNumber": "INV-2024-001",
"purchaseOrder": null,
"total": 1250.00,
"lineItems": [
{ "description": "Consulting", "quantity": 10, "unitPrice": 125.00 }
]
}
}Tip: A null value means the data wasn't found in the document. This is more reliable than a guessed value, so you can flag nulls for manual review if needed.
Reviewing and downloading results
Upload one or more files in the Playground and choose Extract. Processing adapts automatically to the number of documents; no mode switch is needed. The Playground opens completed documents in Results: labeled fields, nested sections and tables for lists such as line items. Expand a nested section to inspect its fields; the first non-empty list opens automatically. Missing values appear as “Not found”. Choose JSON to see the original structured result, or Document text to read and download the converted source as a Markdown (.md) file when it is still available.
- Open Download to choose Excel, CSV or JSON. Copy and remove actions appear on hover or keyboard focus on desktop and remain visible on touch screens.
- Excel (.xlsx) downloads contain a Documents sheet and separate sheets for nested lists.
- CSV downloads produce one .csv file for a single table, or a .zip archive of CSV files when the result has multiple tables.
- Document and row IDs link nested rows to their source. Field paths preserve nested names; blank cells represent missing values.
- Copy fields or Copy table pastes data into your spreadsheet. JSON (.json) downloads preserve the exact structured result, including the original field names and values.
- Select any completed document to review or download it while the remaining batch is still processing. Batch downloads include completed documents only; failed and cancelled jobs are excluded. History has the same views and download options.
Views and downloads use the existing result, with no reprocessing or extra charge. The API returns JSON only; Excel and CSV are Playground download formats.
Webhooks
A webhook sends a signed event to your HTTPS endpoint each time a run of a workflow finishes, so you do not have to poll. Add one on the workflow page under Integrations, then choose Send test to check your receiver. Smole shows the signing secret once, when you add the webhook or rotate its secret; store it in your receiver's configuration straight away.
Events
run.completed- A document finished extracting. The body carries the result unless you turned it off.run.failed- A document could not be processed.resultisnullanderrorsays why, with the failing fields when the result did not match the schema. Failed runs are not charged.ping- Sent by Send test. It never runs a document and is free.
{
"id": "0b6f4b1e-9c1d-4a5e-8f3a-2d7c6e5b4a39",
"type": "run.completed",
"createdAt": "2026-10-01T09:30:12.000Z",
"workspaceId": "…",
"data": {
"runId": "…",
"status": "completed",
"filename": "invoice.pdf",
"processingMode": "normal",
"workflowId": "…",
"workflowVersionId": "…",
"workflowVersion": 3,
"schemaId": "…",
"schemaVersion": 7,
"source": null,
"error": null,
"result": { "invoiceNumber": "INV-2024-001", "total": 1250.0 },
"resultTruncated": false
}
}The event id and the body bytes are the same on every attempt. A result over 256 KiB is left out: result is null, resultTruncated is true, and your receiver reads it with GET /api/pipeline/:runId and an API key.
A run.failed event has the same data fields, with status "failed", result null and an error object:
"error": {
"code": "RESULT_SCHEMA_INVALID",
"message": "Extraction failed: Final extraction does not match the submitted schema",
"stage": "extraction",
"retryable": false,
"errors": [
"$.total is required but the document does not state it",
"$.lines[0].qty must be number"
]
}code- The failure code, ornullwhen none was recorded.message- A short description, at most 500 characters, ornull.stage- Where the run failed, for exampleconversionorextraction, ornull.retryable- Whether submitting the document again can succeed, ornullwhen unknown.errors- ForRESULT_SCHEMA_INVALID, the failing field paths (at most 50). An empty array for every other code.
Headers and verification
Events follow Standard Webhooks, so any Standard Webhooks library can verify them.
webhook-id- The event ID, the same on every attempt.webhook-timestamp- The attempt time in Unix seconds.webhook-signature-v1,<base64>, an HMAC-SHA256 over<id>.<timestamp>.<raw body>. For 24 hours after you rotate the secret it holds two values, one per secret.smole-event-type-run.completed,run.failedorping.
// npm install standardwebhooks
import http from "node:http";
import { Webhook } from "standardwebhooks";
const webhook = new Webhook(process.env.WEBHOOK_SECRET); // whsec_...
const seen = new Set(); // use a durable store in production
http.createServer((request, response) => {
const chunks = [];
request.on("data", (chunk) => chunks.push(chunk));
request.on("end", () => {
let event;
try {
// Verify the raw body before parsing it. Refuses timestamps over 5 minutes old.
event = webhook.verify(Buffer.concat(chunks).toString("utf8"), request.headers);
} catch {
return response.writeHead(400).end();
}
if (!seen.has(event.id)) {
seen.add(event.id);
// handle event.type and event.data here
}
response.writeHead(204).end(); // answer 2xx for a repeat too
});
}).listen(8787);The complete receiver, with tests, is the examples/webhook-receiver project in the Smole repository.
Retries and redelivery
- Any 2xx answer delivers the event. Smole does not read the response body.
- 408, 425, 429, 5xx, a timeout (10 s) or a connection error is retried after about 1 min, 5 min, 15 min, 1 h, 4 h and 12 h, honouring
Retry-Afterup to 12 h. After 7 attempts the delivery fails. - 401 or 403 fails the delivery and marks the webhook as needing attention. 410 fails it and pauses the webhook. Redirects are never followed. Any other 4xx fails the delivery at once.
- After three failed deliveries in a row, Smole pauses the webhook. Events of later runs wait for up to 7 days until you resume it.
- Delivery is at least once and in no guaranteed order. Keep the
webhook-idvalues you processed and answer a repeat with 2xx. - The webhook page lists every delivery with its attempts. Filter by Failed to see the dead letters, and choose Redeliver to send one again with the same ID and body. Redelivery never processes the document again or charges for it.
Webhook URLs must be public HTTPS on port 443, without credentials in the URL. Smole checks the address on every attempt and refuses local and private networks. See API Reference to manage webhooks with an API key.
Google Sheets
A Google Sheet integration adds a row to a spreadsheet each time a run of a workflow completes. Connect a Google account under Account, Connections, then choose Add Google Sheet in the Integrations section of the workflow page. Smole creates a new spreadsheet called SMOLE – <workflow title> in that account's Drive. It cannot write to a spreadsheet you already have.
What Smole can access
- The only Google Drive permission Smole asks for is
drive.file. It can create files and open the files it created. It cannot see, list or open anything else in your Drive. - Smole reads the Google account's email address to label the connection.
- Disconnecting the account stops every write and deletes Smole's stored access. The spreadsheets stay in your Drive, and the account's email address stays in the workspace's integration history.
- Deleting the integration does not delete its spreadsheet.
Layout
Documents- One row per completed run:$runId,$sourceFilename,$completedAt,$workflowVersion, then one column per field.- One tab per list field, named by its path (for example
line_items) - One row per item:$runId,$rowId,$parentRowId,$index, then the item's fields. This is the same layout as the XLSX export. - Values are written as they were extracted: text is never turned into a formula or a date, and numbers stay numbers. Text over 50,000 characters is cut and ends in
…[truncated]. - When a later run has a new field, its column is added on the right. Existing columns are never moved or renamed.
- Only completed runs are written. Failed runs are never written and are not charged.
Do not restructure the Smole tabs
Keep row 1 and column A of every Smole tab as they are: column A holds $runId, which Smole reads before each write so that a run is never written twice. Adding your own columns to the right, moving other columns and adding your own tabs are safe. A Smole tab you delete is created again with its header on the next run.
- If cell
A1of a Smole tab is no longer$runId, Smole stops and marks the integration as needing attention rather than write over your data. Put$runIdback inA1of every Smole tab, then choose Resume. - If the spreadsheet is deleted or reaches Google's size limit, the integration also needs attention. Restore the spreadsheet or free up space, then resume, or add a new Google Sheet.
- If Google no longer accepts the connection, the integration page offers Connect <email> again: choose it and pick that Google account, and writing resumes by itself. If the account was disconnected in Smole, or the integration was paused, connect it again and then choose Resume.
Testing and redelivery
- Send test checks that Smole can open the spreadsheet and that the tab layout is intact. It writes nothing.
- Runs are written one at a time, a few seconds apart, so a large batch takes a while to appear.
- Redeliver never adds a second row for a run that is already in the spreadsheet, and never processes the document again or charges for it.
Batch Processing
Process multiple documents by submitting them in parallel and polling for results:
# Submit multiple documents
for file in invoices/*.pdf; do
curl -s -X POST https://api.smole.tech/api/pipeline/file \
-H "X-API-Key: ak_your_api_key" \
-F "file=@$file" \
-F "schemaId=your_schema_id" >> job_ids.txt
echo "" >> job_ids.txt
done
# Poll each job for completion
# (use jq to extract the job ID from each response)