Schemas

Define the structure of your extracted data.

What is a Schema?

A schema defines the structure of the data you want to extract from a document. It tells the AI exactly which fields to look for, what type each field should be, and how the output should be organized. Schemas follow the JSON Schema standard.

Create a schema once and reuse it across any number of documents. For example, create an "Invoice" schema and use it to extract data from hundreds of different invoices. The same fields get extracted every time.

Schemas and Workflows

A schema says what to extract. A workflow says how to process a kind of document: it references one schema and adds its own preparation settings (force OCR, preserve tables, preserve links, max pages), and it keeps the history of every run. One schema can back several workflows. For example, "Invoices (scans)" forces OCR and "Invoices (digital)" does not, and both use the same invoice schema.

A workflow can use a built-in schema or one of your own. Editing a schema changes it in place: workflows that use it pick up the change on their next run, and every run records the schemaVersion it used. See API Reference for the workflow endpoints.

Schema Structure

Every schema is a JSON object with type, properties, and optionally required:

JSON
{
  "type": "object",
  "properties": {
    "companyName": { "type": "string" },
    "invoiceNumber": { "type": "string" },
    "date": { "type": "string" },
    "total": { "type": "number" },
    "paid": { "type": "boolean" },
    "lineItems": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "description": { "type": "string" },
          "quantity": { "type": "integer" },
          "unitPrice": { "type": "number" }
        }
      }
    }
  },
  "required": ["companyName", "invoiceNumber", "total"]
}

Supported Types

TypeDescriptionExample Value
stringText values: names, dates, IDs, descriptions"INV-2024-001"
numberDecimal numbers: prices, totals, percentages149.99
integerWhole numbers: quantities, counts, years42
booleanTrue/false values: flags, checkboxestrue
arrayLists of items: line items, tags, addresses. Define items for element type["tag1", "tag2"]
objectNested groups: address, contact info. Define properties for fields{"city": "NYC"}

Enum Fields

Use enum to restrict a field to a fixed set of allowed values. This is useful for categories, statuses, or any field with a known set of options:

JSON
{
  "status": {
    "type": "string",
    "enum": ["paid", "unpaid", "overdue", "cancelled"]
  },
  "priority": {
    "type": "string",
    "enum": ["low", "medium", "high"]
  }
}

The AI will only output one of the listed values. If the document's value doesn't clearly match any option, the field will be returned as null.

Nesting and Arrays

Schemas can be nested to any depth. Use object for grouped fields and array for repeated structures:

JSON
{
  "type": "object",
  "properties": {
    "vendor": {
      "type": "object",
      "properties": {
        "name": { "type": "string" },
        "address": {
          "type": "object",
          "properties": {
            "street": { "type": "string" },
            "city": { "type": "string" },
            "zip": { "type": "string" }
          }
        }
      }
    },
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "qty": { "type": "integer" },
          "price": { "type": "number" }
        }
      }
    }
  }
}

Arrays can be empty ([]) if no matching items are found in the document. Each object inside an array follows the same field rules: fields not found in the document are returned as null.

Missing Data and Null Values

Important: When the AI cannot find a field's value in the document, it returns null instead of guessing. This is by design: you get accurate data or an explicit signal that the data wasn't found.

Any optional leaf field (string, number, integer, boolean) can be null in the extraction result. This ensures you never receive fabricated or hallucinated data.

Example

Given this schema and a document that only contains a company name and total:

JSON
// Schema
{
  "type": "object",
  "properties": {
    "companyName": { "type": "string" },
    "invoiceNumber": { "type": "string" },
    "purchaseOrder": { "type": "string" },
    "total": { "type": "number" }
  }
}

// Result: fields not found in the document are null
{
  "companyName": "Acme Corp",
  "invoiceNumber": null,
  "purchaseOrder": null,
  "total": 1250.00
}

This applies to all field types:

TypeWhen FoundWhen Not Found
string"Acme Corp"null
number149.99null
integer5null
booleantruenull
enum"paid"null
array[item1, item2][] (empty array)

Tip: When processing results, always check for null before using a field's value. This pattern gives you reliable data. If a value is present, you can trust it came from the document.

Required Fields and Value Constraints

If the document lacks a required field, or an extracted value breaks a schema keyword such as pattern, minimum or maxLength, the extraction fails with RESULT_SCHEMA_INVALID. The job's error lists the offending paths, such as /invoiceNumber. errorRetryable is false, and the failed job is not charged. Smole never invents or alters a value to make it fit.

Creating and Managing Schemas

Create schemas via the API and reuse them across multiple documents. See API Reference for all schema endpoints.

Terminal
# Create a schema
curl -X POST https://api.smole.tech/api/schemas \
  -H "X-API-Key: ak_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Receipt Schema",
    "description": "Extract data from retail receipts",
    "jsonSchema": {
      "type": "object",
      "properties": {
        "store": { "type": "string" },
        "items": {
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "name": { "type": "string" },
              "price": { "type": "number" }
            }
          }
        },
        "total": { "type": "number" }
      }
    }
  }'

# List your schemas
curl https://api.smole.tech/api/schemas \
  -H "X-API-Key: ak_your_api_key"

# Update a schema
curl -X PUT https://api.smole.tech/api/schemas/{schema_id} \
  -H "X-API-Key: ak_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{ "name": "Updated Receipt Schema" }'

# Delete a schema
curl -X DELETE https://api.smole.tech/api/schemas/{schema_id} \
  -H "X-API-Key: ak_your_api_key"

Updating the jsonSchema field automatically increments the schema version. Previous extraction results are not affected.

A schema that an active workflow uses cannot be deleted: the request returns 409 SCHEMA_IN_USE. Archive the workflow first, or point it at another schema.

AI-Generated Schemas

Don't want to write JSON Schema by hand? Use the /api/schemas/generate endpoint to have AI create a schema based on your field definitions:

Terminal
curl -X POST https://api.smole.tech/api/schemas/generate \
  -H "X-API-Key: ak_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Invoice Schema",
    "documentType": "invoice",
    "fields": [
      {
        "name": "invoiceNumber",
        "description": "Unique invoice identifier",
        "type": "string",
        "required": true
      },
      {
        "name": "vendor",
        "description": "Company that issued the invoice",
        "type": "string",
        "required": true
      },
      {
        "name": "lineItems",
        "description": "List of items on the invoice",
        "type": "array",
        "required": true,
        "itemType": "object",
        "nestedFields": [
          { "name": "description", "description": "Item description", "type": "string", "required": true },
          { "name": "quantity", "description": "Number of units", "type": "number", "required": true },
          { "name": "unitPrice", "description": "Price per unit", "type": "number", "required": true, "format": "currency" }
        ]
      },
      {
        "name": "total",
        "description": "Total amount due",
        "type": "number",
        "required": true,
        "format": "currency"
      }
    ]
  }'

The schema is generated asynchronously. Poll GET /api/schemas/:id until the status is ready.

Field Definition Options

PropertyRequiredDescription
nameYesField name in the output JSON
descriptionYesDescription to help AI understand what to extract
typeYesstring, number, boolean, date, array, or object
requiredYesWhether this field must be present in the output
formatNoemail, phone, currency, percentage, url, date
itemTypeNoFor arrays: type of items (string, number, or object)
nestedFieldsNoFor objects/arrays of objects: nested field definitions
examplesNoExample values to help AI understand the expected format

Document Types

Specify documentType to optimize schema generation:

invoicereceiptcontractresumereportformother

You can also provide industry (healthcare, finance, legal, retail, general) and exampleMarkdown for even better results.

Schema Best Practices

Use descriptive property names

Name fields to match how they appear in the document. invoiceNumber is better than id because it helps the AI find the right value.

Add description fields

Use description to guide the AI when field names alone are ambiguous. For example: "description": "The date the invoice was issued, not the due date".

Use specific types

Use number for amounts, integer for counts, boolean for yes/no fields. Avoid using string for everything; typed fields give you cleaner data.

Define array item schemas

Always define the items schema for arrays. Without it, the AI has no guidance on the structure of each item.

Handle null in your code

Any field can be null if the data isn't found in the document. Design your data pipeline to handle nulls gracefully, for example with fallback values or by flagging records that need review.

Keep schemas focused

Extract only the fields you need. A schema with 5 relevant fields will produce better results than one with 50 fields where most are irrelevant to the document.

Valid vs Invalid Schemas

The root of every schema must be "type": "object" with a properties field. Here are common mistakes:

Invalid: missing root object

JSON
{
  "invoiceNumber": { "type": "string" },
  "total": { "type": "number" }
}

Properties must be nested inside "properties" with "type": "object" at the root.

Valid: correct structure

JSON
{
  "type": "object",
  "properties": {
    "invoiceNumber": { "type": "string" },
    "total": { "type": "number" }
  }
}

Invalid: array items missing schema

JSON
{
  "type": "object",
  "properties": {
    "items": { "type": "array" }
  }
}

Arrays need an items property defining the element type.

Valid: array with item schema

JSON
{
  "type": "object",
  "properties": {
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "price": { "type": "number" }
        }
      }
    }
  }
}

Invalid: unsupported type

JSON
{
  "type": "object",
  "properties": {
    "date": { "type": "date" }
  }
}

JSON Schema has no date type. Use "type": "string". The AI will format dates as ISO 8601 (YYYY-MM-DD).