Schemas
Define the structure of your extracted data.
What is a Schema?
A schema defines the structure of the data you want to extract from a document. It tells the AI exactly which fields to look for, what type each field should be, and how the output should be organized. Schemas follow the JSON Schema standard.
Create a schema once and reuse it across any number of documents. For example, create an "Invoice" schema and use it to extract data from hundreds of different invoices. The same fields get extracted every time.
Schemas and Workflows
A schema says what to extract. A workflow says how to process a kind of document: it references one schema and adds its own preparation settings (force OCR, preserve tables, preserve links, max pages), and it keeps the history of every run. One schema can back several workflows. For example, "Invoices (scans)" forces OCR and "Invoices (digital)" does not, and both use the same invoice schema.
A workflow can use a built-in schema or one of your own. Editing a schema changes it in place: workflows that use it pick up the change on their next run, and every run records the schemaVersion it used. See API Reference for the workflow endpoints.
Schema Structure
Every schema is a JSON object with type, properties, and optionally required:
{
"type": "object",
"properties": {
"companyName": { "type": "string" },
"invoiceNumber": { "type": "string" },
"date": { "type": "string" },
"total": { "type": "number" },
"paid": { "type": "boolean" },
"lineItems": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": { "type": "string" },
"quantity": { "type": "integer" },
"unitPrice": { "type": "number" }
}
}
}
},
"required": ["companyName", "invoiceNumber", "total"]
}Supported Types
Enum Fields
Use enum to restrict a field to a fixed set of allowed values. This is useful for categories, statuses, or any field with a known set of options:
{
"status": {
"type": "string",
"enum": ["paid", "unpaid", "overdue", "cancelled"]
},
"priority": {
"type": "string",
"enum": ["low", "medium", "high"]
}
}The AI will only output one of the listed values. If the document's value doesn't clearly match any option, the field will be returned as null.
Nesting and Arrays
Schemas can be nested to any depth. Use object for grouped fields and array for repeated structures:
{
"type": "object",
"properties": {
"vendor": {
"type": "object",
"properties": {
"name": { "type": "string" },
"address": {
"type": "object",
"properties": {
"street": { "type": "string" },
"city": { "type": "string" },
"zip": { "type": "string" }
}
}
}
},
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": { "type": "string" },
"qty": { "type": "integer" },
"price": { "type": "number" }
}
}
}
}
}Arrays can be empty ([]) if no matching items are found in the document. Each object inside an array follows the same field rules: fields not found in the document are returned as null.
Missing Data and Null Values
Important: When the AI cannot find a field's value in the document, it returns null instead of guessing. This is by design: you get accurate data or an explicit signal that the data wasn't found.
Any optional leaf field (string, number, integer, boolean) can be null in the extraction result. This ensures you never receive fabricated or hallucinated data.
Example
Given this schema and a document that only contains a company name and total:
// Schema
{
"type": "object",
"properties": {
"companyName": { "type": "string" },
"invoiceNumber": { "type": "string" },
"purchaseOrder": { "type": "string" },
"total": { "type": "number" }
}
}
// Result: fields not found in the document are null
{
"companyName": "Acme Corp",
"invoiceNumber": null,
"purchaseOrder": null,
"total": 1250.00
}This applies to all field types:
Tip: When processing results, always check for null before using a field's value. This pattern gives you reliable data. If a value is present, you can trust it came from the document.
Required Fields and Value Constraints
If the document lacks a required field, or an extracted value breaks a schema keyword such as pattern, minimum or maxLength, the extraction fails with RESULT_SCHEMA_INVALID. The job's error lists the offending paths, such as /invoiceNumber. errorRetryable is false, and the failed job is not charged. Smole never invents or alters a value to make it fit.
Creating and Managing Schemas
Create schemas via the API and reuse them across multiple documents. See API Reference for all schema endpoints.
# Create a schema
curl -X POST https://api.smole.tech/api/schemas \
-H "X-API-Key: ak_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"name": "Receipt Schema",
"description": "Extract data from retail receipts",
"jsonSchema": {
"type": "object",
"properties": {
"store": { "type": "string" },
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": { "type": "string" },
"price": { "type": "number" }
}
}
},
"total": { "type": "number" }
}
}
}'
# List your schemas
curl https://api.smole.tech/api/schemas \
-H "X-API-Key: ak_your_api_key"
# Update a schema
curl -X PUT https://api.smole.tech/api/schemas/{schema_id} \
-H "X-API-Key: ak_your_api_key" \
-H "Content-Type: application/json" \
-d '{ "name": "Updated Receipt Schema" }'
# Delete a schema
curl -X DELETE https://api.smole.tech/api/schemas/{schema_id} \
-H "X-API-Key: ak_your_api_key"Updating the jsonSchema field automatically increments the schema version. Previous extraction results are not affected.
A schema that an active workflow uses cannot be deleted: the request returns 409 SCHEMA_IN_USE. Archive the workflow first, or point it at another schema.
AI-Generated Schemas
Don't want to write JSON Schema by hand? Use the /api/schemas/generate endpoint to have AI create a schema based on your field definitions:
curl -X POST https://api.smole.tech/api/schemas/generate \
-H "X-API-Key: ak_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"name": "Invoice Schema",
"documentType": "invoice",
"fields": [
{
"name": "invoiceNumber",
"description": "Unique invoice identifier",
"type": "string",
"required": true
},
{
"name": "vendor",
"description": "Company that issued the invoice",
"type": "string",
"required": true
},
{
"name": "lineItems",
"description": "List of items on the invoice",
"type": "array",
"required": true,
"itemType": "object",
"nestedFields": [
{ "name": "description", "description": "Item description", "type": "string", "required": true },
{ "name": "quantity", "description": "Number of units", "type": "number", "required": true },
{ "name": "unitPrice", "description": "Price per unit", "type": "number", "required": true, "format": "currency" }
]
},
{
"name": "total",
"description": "Total amount due",
"type": "number",
"required": true,
"format": "currency"
}
]
}'The schema is generated asynchronously. Poll GET /api/schemas/:id until the status is ready.
Field Definition Options
Document Types
Specify documentType to optimize schema generation:
You can also provide industry (healthcare, finance, legal, retail, general) and exampleMarkdown for even better results.
Schema Best Practices
Use descriptive property names
Name fields to match how they appear in the document. invoiceNumber is better than id because it helps the AI find the right value.
Add description fields
Use description to guide the AI when field names alone are ambiguous. For example: "description": "The date the invoice was issued, not the due date".
Use specific types
Use number for amounts, integer for counts, boolean for yes/no fields. Avoid using string for everything; typed fields give you cleaner data.
Define array item schemas
Always define the items schema for arrays. Without it, the AI has no guidance on the structure of each item.
Handle null in your code
Any field can be null if the data isn't found in the document. Design your data pipeline to handle nulls gracefully, for example with fallback values or by flagging records that need review.
Keep schemas focused
Extract only the fields you need. A schema with 5 relevant fields will produce better results than one with 50 fields where most are irrelevant to the document.
Valid vs Invalid Schemas
The root of every schema must be "type": "object" with a properties field. Here are common mistakes:
Invalid: missing root object
{
"invoiceNumber": { "type": "string" },
"total": { "type": "number" }
}Properties must be nested inside "properties" with "type": "object" at the root.
Valid: correct structure
{
"type": "object",
"properties": {
"invoiceNumber": { "type": "string" },
"total": { "type": "number" }
}
}Invalid: array items missing schema
{
"type": "object",
"properties": {
"items": { "type": "array" }
}
}Arrays need an items property defining the element type.
Valid: array with item schema
{
"type": "object",
"properties": {
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": { "type": "string" },
"price": { "type": "number" }
}
}
}
}
}Invalid: unsupported type
{
"type": "object",
"properties": {
"date": { "type": "date" }
}
}JSON Schema has no date type. Use "type": "string". The AI will format dates as ISO 8601 (YYYY-MM-DD).