# 0sql Semantic Modeling and Query Reference > Complete reference for AI agents building 0sql semantic models and query specs 0sql is a semantic layer as a service. Developers describe warehouse tables, dimensions, measures, joins and row-level security policies in YAML, deploy the project with the `zsql` CLI, and POST query specs to get back the SQL for their warehouse. 0sql plans SQL; it never executes it. ## 0sql Semantic Modeling Rules (Must Follow) These rules are mandatory. AI agents generating 0sql YAML or query specs must follow them exactly. ### Naming Rules 1. A field name identifies ONE concept across the entire semantic layer. The same name on several tables is merged into that one field (the planner picks the table by requested dimensions, then lowest cost). Never reuse a name for a different concept. Names are scoped per type: the same name may exist as both a dimension and a measure (e.g. "Days To Recontact") 2. Field names MUST be stable - a rename changes the field's uid and breaks every spec your applications already send; there is no migration mechanism 3. Table names MUST be unique within a datasource 4. Use descriptive, business-friendly names (not database column names) 5. Field names should use Title Case with spaces (e.g., "Total Revenue", "Customer Name") ### Structure Rules 1. Every table MUST define: datasource, name, physical_name, cost, fields 2. Every field MUST define: type, name, data_type, expression 3. Relations are directional and defined separately from tables (in rel.*.yml files) 4. YAML is declarative - key order does not matter, but meaning does 5. One table per tbl.*.yml file 6. Multiple relationships can be in one rel.*.yml file (grouped by domain) ### Relationship Rules 1. many_to_many is NOT supported - use junction/bridge tables instead 2. Cardinality must match actual data: many_to_one, one_to_many, one_to_one 3. allow_measure_expansion should be used carefully (can cause double-counting) 4. Join conditions use `left.column = right.column` syntax 5. The "left" table is the table containing the foreign key 6. Relations are defined on the child table pointing to the parent ### Expression Rules 1. A measure whose expression reads table columns MUST include an aggregation function (sum, count, avg, min, max) 2. Dimensions should NOT include aggregation functions 3. Field references use bracket notation: [Field Name]; add @m or @d ([Net Paid]@m, [Category]@d) when a dimension and a measure share the name 4. A compound measure references other fields in brackets, never raw columns; a formula built only from measure references needs no aggregate of its own 5. SQL expressions are adapter-specific (PostgreSQL, Snowflake, etc.) 6. Use `primary_key: true` on the unique identifier dimension ### Cost Rules 1. Costs are relative indicators for query routing, not pricing 2. Lower cost = preferred table when multiple tables can answer a query 3. Dimension tables: typically cost 10 4. Fact tables: typically cost 100 5. Hot tier (fast storage): lower cost 6. Cold tier (archival): higher cost ### File Organization Rules 1. Tables: `tbl.{name}.yml` (e.g., tbl.orders.yml, tbl.customers.yml) 2. Relations: `rel.{domain}.yml` (e.g., rel.sales.yml, rel.inventory.yml) 3. One relation file per domain grouping related joins 4. Do NOT duplicate physical tables across semantic tables 5. Organize models by domain in subdirectories (models/sales/, models/inventory/) ### Data Type Rules 1. Supported data types: string, integer, bigint, decimal, date, date_time, boolean, binary 2. Use decimal for monetary values 3. Use date for dates without time, date_time for timestamps 4. Use bigint for large integer values (counts, IDs) ### Deploy Rules 1. `zsql check` validates the project without deploying; `zsql deploy` ships the checked-out git branch to the service 2. `production_branch` in project.yml names the branch applications read from by default 3. Secrets never go in YAML: the API key lives in `.zsql`, and `zsql deploy` strips datasource secrets from `datasources.yml` before building the archive 4. `tests/*.yml` run on every deploy; a failing test makes `zsql deploy` exit non-zero ### Query Spec Rules 1. Projections reference fields by name (case-insensitive), uid or synonym, never by warehouse column 2. Add `@m` or `@d` to a name when a dimension and a measure share it (`"field": "Date@d"`) 3. `filters` is a flat array (AND) or one `{"and": [...]}` / `{"or": [...]}` tree; a filter on a measure becomes HAVING 4. Calculations reference measures as `[Name]@m`, dimensions as `[Name]@d`, other projections as `[Alias]`; a calculation with an aggregate function is a measure calculation 5. Never write a bare column in a calculation; the error is `Sql columns like x,y are not permitted in calculations.` 6. Every segment needs `keys` (the dimensions that identify a member); `apply_to` restricts it to named measure projections, otherwise it constrains the whole query 7. Decorators go on the projection: `{"type": "truncate", "grain": "month"}` on dates; `temporalize`, `window`, `contribute` on measures 8. Send a `context` whenever the branch has security policies; the service answers `400 ContextRequired` otherwise ## Table of Contents - [Getting Started](#getting-started) - [Core Concepts](#core-concepts) - [Key Terminologies](#key-terminologies) - [Installing zsql](#installing-zsql) - [Quickstart](#quickstart) - [Tutorial: TPC-DS with DuckDB](#tutorial-tpc-ds-with-duckdb) - [Query API](#query-api) - [Calculations](#calculations) - [Security context](#security-context) - [Discovery: explore, fields and tables](#discovery-explore-fields-and-tables) - [Errors and corrections](#errors-and-corrections) - [Explain](#explain) - [Filters](#filters) - [Projections and decorators](#projections-and-decorators) - [Segments](#segments) - [Shorthand expressions](#shorthand-expressions) - [The query spec](#the-query-spec) - [CLI Reference](#cli-reference) - [Authentication and configuration](#authentication-and-configuration) - [CI/CD](#cicd) - [Datasources](#datasources) - [Deploying](#deploying) - [Modeling with zsql](#modeling-with-zsql) - [Projects](#projects) - [Querying from the CLI](#querying-from-the-cli) - [Command reference](#command-reference) - [Tests](#tests) - [Semantic Model](#semantic-model) - [Athena Adapter](#athena-adapter) - [BigQuery Adapter](#bigquery-adapter) - [ClickHouse Adapter](#clickhouse-adapter) - [Databricks Adapter](#databricks-adapter) - [Druid Adapter](#druid-adapter) - [DuckDB Adapter](#duckdb-adapter) - [MySQL Adapter](#mysql-adapter) - [PostgreSQL Adapter](#postgresql-adapter) - [Redshift Adapter](#redshift-adapter) - [Snowflake Adapter](#snowflake-adapter) - [SQLite Adapter](#sqlite-adapter) - [SQL Server Adapter](#sql-server-adapter) - [Trino Adapter](#trino-adapter) - [Datasources](#datasources) - [Expressions](#expressions) - [Array Expressions](#array-expressions) - [Lookup Expressions](#lookup-expressions) - [SQL Expressions](#sql-expressions) - [Extended Blending Groups](#extended-blending-groups) - [Fields and Types](#fields-and-types) - [Data Types](#data-types) - [Date and DateTime Dimensions](#date-and-datetime-dimensions) - [Field format](#field-format) - [Compound Measures](#compound-measures) - [Snapshot Measures](#snapshot-measures) - [Imports](#imports) - [Partitions](#partitions) - [Cardinality](#cardinality) - [Join Types](#join-types) - [Tables](#tables) - [Advanced Features](#advanced-features) - [Cost Optimization](#cost-optimization) - [Exclusions](#exclusions) - [Inclusions](#inclusions) - [Semantic Routing](#semantic-routing) - [Universe Formation](#universe-formation) - [Security](#security) - [Accounts and keys](#accounts-and-keys) - [Examples](#examples) - [Fact-Dimension Pattern](#fact-dimension-pattern) - [Snowflake Schema Pattern](#snowflake-schema-pattern) - [Star Schema Pattern](#star-schema-pattern) - [Customer 360 Recipe](#customer-360-recipe) - [Sales Analysis Recipe](#sales-analysis-recipe) - [Cohorts and segments](#cohorts-and-segments) - [Complex measures](#complex-measures) - [Cross-fact blend](#cross-fact-blend) - [Filters and top-n](#filters-and-top-n) - [Level of detail](#level-of-detail) - [Period over period](#period-over-period) - [Security context](#security-context) - [Share and windows](#share-and-windows) - [TPC-DS Tutorial](#tpc-ds-tutorial) - [Troubleshooting](#troubleshooting) - [Canonical YAML Examples](#canonical-yaml-examples) - [Common Mistakes to Avoid](#common-mistakes-to-avoid) --- ## Getting Started ### Core Concepts The fundamental concepts of 0sql's semantic layer: tables, relationships, datasources, the query spec, and how a spec becomes one SQL statement. ## What is a Semantic Layer? A semantic layer is an abstraction that sits between your raw database tables and the applications that ask questions of them. It translates business-friendly field names (like "Total Revenue") into one SQL statement against your physical database. **Benefits:** - **One request shape**: Applications ask for fields by name and never write joins - **Consistency**: Single source of truth for measures and definitions - **Routing**: The planner picks the cheapest table and datasource that can answer - **Security**: Row-level policies are applied inside the generated SQL **info: Git-Based Workflow** **All semantic models are version-controlled YAML files.** Unlike UI-based BI tools where changes are made through point-and-click interfaces: - **Version control**: Full Git history for all changes - **Code review**: PR-based workflows for model changes - **Branching**: Test changes in branches before deploying to production - **CI/CD**: Automated testing and deployment pipelines ## Key Components ### Tables Tables (`tbl.*.yml`) represent physical database tables with semantic metadata. They define: - **Dimensions**: Fields to group data or generally provides descriptive information (e.g., Customer Name, Product Category) - **Measures**: Aggregatable quantities, typically numeric (e.g., Total Sales, Average Order Value) - **Cost**: Query optimization hint - **Partitions**: Data availability constraints **Example:** ```yaml name: Store Sales physical_name: store_sales datasource: tpcds cost: 100 # Partition logic helps the engine choose alternative tables when this one # can't satisfy the query. A partition is a predicate on a dimension: a date # or number range, or in_list on any dimension. partitions: - dimension: Date predicate: between filter_value: 24m # 24 months ago filter_value_end: 1d # yesterday fields: - type: dimension name: Store Name data_type: string expression: sql: s_store_name - type: measure name: Total Revenue data_type: decimal expression: sql: sum(ss_sales_price) ``` ### Relationships Relations (`rel.*.yml`) define how tables join together. They specify: - **Left and right tables**: Which tables to join - **Join condition**: SQL expression for the join - **Cardinality**: Relationship type (one-to-one, one-to-many, many-to-one) **Example:** ```yaml datasource: tpcds store_sales_store: left: Store Sales right: Store sql: left.ss_store_sk = right.s_store_sk cardinality: many_to_one ``` ### Datasources Datasources (`datasources.yml`) describe the warehouses your tables live in. 0sql never connects to them; the planner needs two things: - **Adapter**: the SQL dialect to write. One of `athena`, `bigquery`, `clickhouse`, `databricks`, `druid`, `duckdb`, `mysql`, `postgres`, `redshift`, `snowflake`, `sqlite`, `sqlserver`, `trino` - **Tier**: `hot`, `warm` or `cold`. When several datasources can answer a spec, the planner prefers the hotter one Connection details may sit beside them as metadata for your own application; 0sql never reads them, and secrets are stripped from every deploy. ### Segments A **segment** is a population: a set of members and the criteria that decide who is in it. "Customers who spent over 10,000 last year", "stores opened after 2020", "sessions that reached checkout". Segments are not modeled in YAML. They are part of the query spec your application sends, under the `segments` key, so nobody changes the model to ask a cohort question. A segment is made of two things: - **Keys**: the dimensions that identify a member (for instance `Customer ID`). They set the grain of the population and are what the planner joins on. Every segment needs at least one. - **Criteria**: `filters` in the same shape as the query's filters (dimension filters, measure conditions, a top-n ranking), and/or `measures` whose fact table defines membership. A segment constrains the **whole query**, or with `apply_to` only the **named measure projections**, so a cohort can sit beside its baseline in one statement: new-customer revenue next to total revenue, without a correlated subquery. `mode` is `include` (keep members, the default) or `exclude` (drop them, an anti-join). ```json { "name": "High spenders", "keys": ["Customer ID"], "filters": [{"field": "Total Revenue", "predicate": "greater_than", "value": "10000"}], "mode": "include", "apply_to": ["Total Revenue"] } ``` The planner emits each segment once as a CTE and joins every consuming node to it on the keys, so a segmented measure keeps the grain safety and row-level security of the model: policies apply inside the segment's CTE too. Populations compose by intersection: "top 5 products" plus a segment means the global top 5 intersected with the segment. **Expanding segments** answer "same X" questions such as *what else is bought in the same basket as sub-commodity A?* Besides its keys, an expanding segment carries a companion dimension of the other members sharing the key into the query, so each row multiplies into one row per companion value. Pair them with distinct counts of the key, since sums would overcount the fan-out. See [Segments](#segments). ### Calculations A **calculation** is a derived column your application adds to a spec: a rate, a ratio, a share, a margin, a difference, a bucket. It is a SQL expression written over the fields of the semantic layer, not over warehouse columns, so it inherits every definition, join and security rule the model already enforces. Calculations are not modeled in YAML; they live under the `calculations` key of the spec. Calculations reference fields by name: - `[Total Revenue]@m`: a measure - `[Item Category]@d`: a dimension - `[Margin]`: another column of the same query, including other calculations The formula decides what kind of column it is. With an aggregate function it is a **measure calculation**, computed over the aggregated result: ```sql sum([Store Quantity]@m) / nullif(sum([Store Quantity]@m) + sum([Inventory Quantity On Hand]@m), 0) ``` Without one it is a **dimension calculation**, evaluated per row before grouping: ```sql case when [Order Total]@d >= 500 then 'Large' when [Order Total]@d >= 100 then 'Medium' else 'Small' end ``` Raw column and table names are rejected (`columns like x are not permitted in calculations`), a formula is an expression and never a `SELECT`, and every reference must resolve to a field of the stated type. In the spec: ```json { "alias": "Sell Through Rate", "sql": "sum([Store Quantity]@m) / nullif(sum([Store Quantity]@m) + sum([Inventory Quantity On Hand]@m), 0)", "data_type": "decimal" } ``` **Calculations and compound measures.** They look alike but sit at different layers. A [compound measure](#compound-measures) is modeled in YAML, deployed with the project and available to every caller. A calculation lives in one request and needs no deployment. A calculation your applications keep sending is a signal to promote it into the model. See [Calculations](#calculations). ### Decorators A **decorator** is a transform on a projection in the spec, with no formula and no model change. Put one on a measure and the planner computes a new column from it as a final pass over the aggregated result: last year's value, its share of the total, its 7-period moving average. The measure's definition, joins, grain rules and row-level security are unchanged; the decorator only transforms the aggregate. Three families cover measures: - **Time transforms** (`temporalize`): year-over-year, quarter-over-quarter, month-over-month, week-over-week, day-over-day, as the prior value or as percent change. The planner reads the query's date dimension and its grain to work out the comparison period, so a month-grain query with a year-over-year transform compares each month to the same month last year. - **Percent of total** (`contribute`): a measure divided by its total, partitioned by the dimensions you name in `partition_refs`. - **Windows** (`window`): moving and running `avg`, `sum`, `min`, `max`, `count`, plus `rank`, `dense_rank`, `percent_rank`, `ntile`, `lag`, `lead`, `first_value`, `last_value`. Windows order by the query's date dimension automatically. Dimensions have their own: `truncate` a date to week, month, quarter or year, or `extract` a part of it (day of week, month name, year-month). ```json {"field": "Total Revenue", "alias": "Revenue LY", "decorators": [{"type": "temporalize", "transform": "year_over_year"}]} ``` A decorator is the fast path for a transform that has a name; a calculation is the general path for a formula that doesn't. See [Projections and decorators](#projections-and-decorators). ## Core Design Principles **tip: One Name, One Concept: Design Philosophy** **A field name identifies exactly one concept of its type across your entire semantic layer.** There are no table prefixes. Names are scoped per type, so "Days To Recontact" can exist as both a dimension and a measure. Defining the same name on several tables is allowed and is how conformed fields work: every definition of "Total Revenue" is treated as the same measure, and any table that carries it can answer a query for it. The planner chooses the table by the dimensions the query requests, then by the table's `cost`, which you can tune to influence tie breaks. Because a shared name means "the same thing", concepts that differ must be named differently. "Caller Country" and "Ship Country" stay separate; a single "Country" on both tables would be merged. **Benefits:** - **Unambiguous references**: When a query asks for "Total Revenue", there's exactly one meaning, whichever table serves it - **Simpler compound measures**: Reference fields by name without table prefixes: `[Total Revenue] - [Total Cost]` - **Blending by naming**: A shared "Date" or "Customer" across fact tables is what makes cross-fact queries possible - **AI-friendly**: LLMs can understand your model without disambiguation **danger: No Many-to-Many Relationships** **0sql does not support many-to-many (`many_to_many`) relationships.** This is an intentional design constraint to ensure predictable query behavior. **Why this limitation exists:** - Prevents double-counting in aggregations (a common source of bugs) - Forces explicit modeling that makes data relationships clearer - Ensures consistent query results **Workaround:** If you need many-to-many relationships (e.g., Users ↔ Roles), create a **junction/bridge table** with two separate relationships: - `Users → UserRoles` (one_to_many) - `UserRoles → Roles` (many_to_one) **Trade-off:** This requires additional YAML files and explicit modeling, but results in more maintainable and predictable semantic models. ### Fact table grain and join safety Every fact table in your model has a **grain**: the level of detail at which each row is stored. 0sql's universe formation and routing logic assume that: - Joins between tables respect their true cardinality (see [Cardinality](#cardinality)). - Many-to-many joins are modeled explicitly via bridge tables rather than hidden in a single relationship. If two facts live at incompatible grains (for example, one is “daily inventory” and the other is “per‑transaction sales”), 0sql will: - **Restrict measures to safe join paths** when building universes: measures are only exposed on paths where every join is one-to-one (or explicitly allows measure expansion). On many-to-one or one-to-many paths, only dimensions are exposed, so measures are not double-counted. - Use **automatic data blending** across universes when no single path can safely serve all measures: each fact is aggregated in its own universe and the results are merged on common dimensions, in one statement. ## How It Works 1. **Model in YAML.** Tables, relationships, datasources and security policies live in a git repository. 2. **Deploy.** `zsql deploy` sends the project to the service under the checked-out git branch. The service validates it, forms universes (every join route from every table) and runs `tests/*.yml`. 3. **Request.** Your application POSTs a spec (projections, filters, calculations, segments) with a security context to `POST /projects/{uid}/branches/{branch}/sql`, or one line of shorthand as `expr`. 4. **Plan.** The planner resolves field names, picks the universe that reaches every field, chooses join routes by cost, tier and partitions, applies the branch's policies, and writes one SQL statement in the datasource's dialect. Planning takes microseconds. 5. **Run.** Your application runs the SQL against its own warehouse. 0sql never executes or stores anything. ```sh zsql sql --expr "item category, store net paid" --context user.json ``` The response names the datasource and holds the statement. See the [Query API](#query-api) for a full request and the SQL that comes back. ## Project Structure A typical 0sql project: ``` my-project/ ├── project.yml # name, uid, production_branch ├── datasources.yml # one warehouse per key: name, adapter, tier ├── security.yml # row-level security policies ├── .zsql # API key and server (gitignored) ├── models/ # tbl.*.yml and rel.*.yml, any folder depth │ ├── sales/ │ │ ├── tbl.orders.yml │ │ ├── tbl.customers.yml │ │ └── rel.sales.yml │ └── inventory/ │ └── tbl.products.yml └── tests/ # planner assertions, run on every deploy └── revenue_positive.yml ``` ## Workflow 1. **Initialize**: `zsql init my-project` 2. **Configure**: describe your warehouses in `datasources.yml` 3. **Model**: `zsql new table NAME --datasource DS` and `zsql new relation --datasource DS`, then edit the YAML 4. **Validate**: `zsql check`, plus `tests/*.yml` 5. **Deploy**: `zsql auth --api-key zsk_... --server https://app.0sql.io` once, then `zsql deploy` 6. **Query**: `zsql sql --expr "..."` to see the SQL, then POST specs from your application ## Next Steps - [Key terms and when to use what](#key-terminologies): glossary and decision guide - [Query API](#query-api): the request, the context, and the SQL that comes back - Learn how to [create tables and relationships](#modeling-with-zsql) - Understand [field types and expressions](#fields-and-types) - Read about [universe formation](#universe-formation) and [advanced features](#advanced-features) ### Key Terminologies Definitions of 0sql-specific and related terms, plus guidance on choosing the right modeling approach. ## Key Terms These terms are used throughout the docs. If you are new to 0sql (or to semantic layers in general), this is a good reference. **Fact table grain**: The level of detail of a fact table (e.g. one row per order line, per day, per invoice). 0sql assumes safe cardinality; incompatible grains are handled via blending or by restricting measures to safe join paths. **Universe**: The complete, cardinality-safe join tree reachable from a chosen root (fact) table. Conceptually, the "virtual table" formed by joining that fact to all dimension tables it can safely reach. The query planner uses universes for query routing and to decide which dimensions can group a given measure. **Automatic data blending**: 0sql's way of combining measures from multiple fact tables (often at different grains) in one query without you writing the joins. The engine aggregates each fact in its own universe and merges results on common dimensions. When dimensions from different tables should count as the same for that merge, you assign them an **extended blend group**. Industry terms you may hear: *data blending*, *multipass SQL*. **Extended blend group**: A named group (e.g. `activity_date`) that you assign to dimensions so they are treated as one logical dimension during **automatic data blending**. Dimensions in the same group are semantically equivalent for the merge, so the planner can combine queries across tables that share that concept. **Snapshot measure**: A measure that captures a value at a **point in time** (beginning or end of a period) rather than summing over the period. Used for stateful quantities: inventory, balances, membership counts. Not for transactional flows. **Compound measure**: A measure whose expression references other measures and dimensions by name (e.g. `[Total Revenue] - [Total Cost]`). Resolved at query time: each referenced measure is planned on its own fact and the results are blended, so a formula can span fact domains (within one datasource). A referenced dimension is resolved within each measure's own universe, so it may sit anywhere along that fact's join path; it must be reachable from every fact involved. No shared universe is required. See also **auto-leveled compound measure**. **Auto-leveled compound measure**: What a cross-domain compound measure becomes when the query groups by a dimension only some of its components can reach. 0sql excludes that dimension from the components that can't reach it, aggregating them at the coarser grain, and joins the results on the dimensions all components share, so the leveled component repeats as a fixed total across the dimension it can't see (`Tickets per 1000 Orders` by Contact Channel: tickets per channel, orders for the month). See: [Compound measures](#compound-measures) **Complex dimension**: A dimension whose expression references other dimensions (e.g. `CASE WHEN [Status Code] = 'A' THEN 'Active' ...`). Only dimensions can be referenced; not measures. Same universe required. **Inclusion measure**: A measure that temporarily **includes** extra dimensions in an intermediate aggregation to compute correct results (like medians or averages), then re-aggregates at the query grain and drops those dimensions in the final result. This solves the problem of computing accurate statistics when underlying data has uneven counts per group (e.g., computing daily median account view hours from per-title data in a streaming company). *Industry analogs for this pattern:* Tableau's **INCLUDE Level of Detail**, MicroStrategy level metrics. Looker does not support this pattern. See: [Inclusions](#inclusions) **Exclusion**: Advanced control that **removes** dimensions from a measure's grouping or filtering behavior (each can be set independently). Used for "% of total" measures and other cases where you want a measure to ignore certain dimensions' grouping or filters. *Industry analogs:* Tableau's **EXCLUDE Level of Detail** , Power BI DAX filter-context functions (e.g. REMOVEFILTERS, ALLEXCEPT). See: [Exclusions](#exclusions) **Segment**: A **population**: key dimensions that identify its members plus the criteria (dimension filters, measure conditions, top-n) that decide membership. Declared in the query spec under `segments`, never in YAML. Constrains the whole query, or with `apply_to` a **single measure** so a cohort can sit beside its baseline in one statement. `include` keeps members; `exclude` drops them. Planned once as a CTE and joined on the keys, so grain safety and row-level security carry through. See: [Segments](#segments) **Expanding segment**: A segment with an `expanding` list that, besides its keys, carries companion dimensions into the query, so each row multiplies into one row per companion value sharing the key ("Sub Commodity (same Basket ID)"). Used for "same basket / same session" analysis. Applies to the whole query; pair it with distinct counts, since sums would overcount the fan-out. **Calculation**: A derived column your application adds to a spec under `calculations`, as a SQL expression over semantic fields (`[Name]@m`, `[Name]@d`, `[Alias]`). With an aggregate it is a measure calculation; without one, a dimension calculation. Never modeled in YAML, never a raw column, validated when the spec is planned, and planned with the same guarantees as a modeled field. Contrast with a **compound measure**, which is modeled and deployed. See: [Calculations](#calculations) **Decorated measure**: A measure projection that carries a `decorators` entry in the spec: a time comparison (`temporalize`: year-over-year, month-over-month, as value or percent change), percent of total (`contribute`), or a window (`window`: moving average, running total, rank, lag/lead). Never in the model; planned as a final pass over the aggregated result, so the measure's definition and security are unchanged. Date dimensions take `truncate` and `extract` the same way. See: [Projections and decorators](#projections-and-decorators) --- ## When to Use What ### Snapshot measure vs normal aggregation - **Use a snapshot measure** when the measure is **stateful**: inventory on hand, account balance, headcount. Summing across days is meaningless; you want the value at the **boundary** of the period (beginning or ending). - **Use a normal measure** (e.g. `sum(amount)`) for **flows**: revenue, transactions, units sold. These are additive over time. ### Extended blend groups vs remodeling tables - **Use extended blend groups** when the same **logical** dimension (e.g. "activity date") exists in multiple tables with different column names or grains, and you want one query to group or filter by that concept across those tables. No need to change table structure. - **Remodel or pre-aggregate** when you need a **different grain** or a dedicated bridge table to fix cardinality; blend groups do not replace proper relationships or grain alignment. ### Inclusion, exclusion, or separate measure? - **Use an inclusion** when you need **multi-level aggregation**: aggregate at an intermediate grain (e.g., sum per account), then compute a statistic over those results (e.g., median across accounts). Essential for correct medians, averages, and percentiles when underlying data has uneven counts per group. - **Use an exclusion** when one measure should behave differently depending on which dimensions are in the query (e.g., "revenue ignoring category" for a %-of-total denominator, or "revenue only by product"). Exclusions control how dimensions affect grouping and filtering. - **Define a separate measure** when the formula or aggregation is genuinely different (e.g., a distinct measure with its own name and definition). Inclusions and exclusions tune **aggregation behavior** of the same underlying expression; they do not change the core formula. ### Validating your model - `zsql check` validates the project on the service without deploying: syntax, references, joins, policies, tests. - `zsql explain --expr "item category, store net paid"` returns the SQL plus the datasource, the table route and the node tree the planner built, so you can **inspect the generated SQL** (CTEs, joins, grouping) before your application sends the spec. - `zsql explore --expr "store net paid"` lists the dimensions and measures that can still join the query, the quickest way to see that a relationship or blend group took effect. If the SQL does not match what you expect, check relationships, grains, extended blend groups, and exclusion/inclusion settings against the docs. ### Installing zsql `zsql` is one static binary with no runtime to install. The installer puts it on your machine; building from source is the alternative. ## Install **Shell script:** ```sh curl -fsSL https://0sql.io/install.sh | sh ``` Works on macOS and Linux. To update, run it again. **From source:** ```sh # from a checkout of the 0sql source cargo build --release -p zsql-cli ./target/release/zsql --version ``` Needs a Rust toolchain. ## Check it ```sh $ zsql --version zsql 0.1.0 $ zsql --help zsql: semantic layer to SQL (0sql.io) ... ``` `zsql --help` prints the flags of one command; the complete list is the [command reference](#command-reference). ## Next steps - [Quickstart](#quickstart): a project, a deploy and the first SQL in a few minutes - [zsql CLI](#zsql-cli): the typical loop and every command - [Authentication and configuration](#authentication-and-configuration): where the key and the server come from ### Quickstart This page takes you from an empty directory to a working request against the hosted service. You will model three TPC-DS tables, deploy them, and get SQL back from the terminal and from curl. Nothing here connects to a warehouse: 0sql only needs to know what your tables look like. ## 1. Get an API key Sign up at [app.0sql.io](https://app.0sql.io). The console creates an account and shows your first **personal key** once. It starts with `zsk_`. Personal keys deploy projects and run queries; later you will create read-only **query keys** (`zqk_`) for your application. See [accounts and keys](#accounts-keys-and-access). ## 2. Install zsql ```sh curl -fsSL https://0sql.io/install.sh | sh zsql --version ``` `zsql` is a single binary that talks to the service. It generates no SQL itself. See [installation](#installing-zsql) for other options. ## 3. Start a project ```sh zsql init tpcds cd tpcds zsql auth --api-key zsk_… --server https://app.0sql.io ``` `zsql init` writes `project.yml`, `datasources.yml`, `security.yml`, empty `models/` and `tests/` directories, and runs `git init` on branch `main`. `zsql auth` stores the key and the server in `.zsql`, which is gitignored, so later commands need neither flag. The project uid is `tpcds`. It is what the service knows the project as, and it appears in every request URL. ## 4. Describe the warehouse 0sql needs the adapter so it emits the right dialect. Connection details are optional metadata; secrets never leave your machine. ```yaml title="datasources.yml" warehouse: name: Warehouse adapter: postgres tier: hot database: tpcds schema: public ``` ## 5. Model three tables Two dimension tables and one fact. Field names are what your requests will use. ```sh zsql new table "Web Sales" --datasource warehouse --domain sales zsql new table Item --datasource warehouse --domain sales zsql new table Date --datasource warehouse --domain sales zsql new relation --datasource warehouse --domain sales ``` ```yaml title="models/sales/tbl.web_sales.yml" name: Web Sales physical_name: web_sales datasource: warehouse cost: 100 fields: - type: measure name: Web Net Paid data_type: decimal format: currency:2 expression: sql: sum(ws_net_paid) - type: measure name: Web Orders data_type: integer expression: sql: count(distinct ws_order_number) ``` ```yaml title="models/sales/tbl.item.yml" name: Item physical_name: item datasource: warehouse cost: 10 fields: - type: dimension name: Category data_type: string synonyms: [department] expression: sql: i_category - type: dimension name: Product Name data_type: string expression: sql: i_product_name ``` ```yaml title="models/sales/tbl.date.yml" name: Date physical_name: date_dim datasource: warehouse cost: 10 fields: - type: dimension name: Date data_type: date expression: sql: d_date ``` ```yaml title="models/sales/rel.sales.yml" datasource: warehouse web_sales_item: left: Web Sales right: Item sql: left.ws_item_sk = right.i_item_sk cardinality: many_to_one web_sales_date: left: Web Sales right: Date sql: left.ws_sold_date_sk = right.d_date_sk cardinality: many_to_one ``` Joins are declared once, with their cardinality. The planner derives every valid route from them. See [relationships](#cardinality). ## 6. Check and deploy ```sh zsql check ``` `zsql check` validates the project on the service and keeps nothing. Then: ```sh zsql deploy ``` ``` deploy tpcds (2184 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 3 tables, 5 fields, 2 joins, 2 paths, 0 policies no tests deployed (tests/*.yml) ``` The first deploy creates the project in your account. The checked-out git branch, `main`, is the deployed branch. See [deploying](#deploying). ## 7. Get SQL from the terminal The shorthand is one comma-separated line: fields to project, predicates to filter, formulas to calculate. ```sh zsql sql --expr "web net paid, category starts with super" ``` ```sql -- datasource: Warehouse SELECT sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") LIKE 'super%' ``` The planner joined `item` because the filter needed it, and compared the string case-insensitively. Try a date grain: ```sh zsql sql --expr "month(date), web net paid" ``` ```sql -- datasource: Warehouse SELECT DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk GROUP BY DATE_TRUNC('month', T1."d_date")::DATE ``` `zsql explain` shows the node tree and timings behind any line, and `zsql repl` keeps a session open. See [querying from the CLI](#querying-from-the-cli). ## 8. Get SQL over HTTP Your application sends the same thing as JSON. Create a query key in the console, grant it the `tpcds` project, then: **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_…" \ -H "Content-Type: application/json" \ -d '{ "spec": { "projections": [ {"field": "Date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Web Net Paid"} ] } }' ``` **JavaScript:** ```js const res = await fetch('https://app.0sql.io/projects/tpcds/branches/main/sql', { method: 'POST', headers: { Authorization: `Bearer ${process.env.ZSQL_QUERY_KEY}`, 'Content-Type': 'application/json' }, body: JSON.stringify({ spec: { projections: [ { field: 'Date', decorators: [{ type: 'truncate', grain: 'month' }] }, { field: 'Web Net Paid' }, ], }, }), }); const { sql, adapter } = await res.json(); // run `sql` against your warehouse with your own client ``` **Python:** ```python r = requests.post( "https://app.0sql.io/projects/tpcds/branches/main/sql", headers={"Authorization": f"Bearer {os.environ['ZSQL_QUERY_KEY']}"}, json={"spec": {"projections": [ {"field": "Date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Web Net Paid"}, ]}}, ) sql = r.json()["sql"] # run `sql` against your warehouse with your own client ``` **expr:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_…" \ -H "Content-Type: application/json" \ -d '{"expr": "month(date), web net paid"}' ``` With `expr` the answer also carries the `spec` the line was read as. ```json { "sql": "SELECT\n\tDATE_TRUNC('month', T1.\"d_date\")::DATE AS \"Month(Date)\",\n\tsum(T0.\"ws_net_paid\") AS \"Web Net Paid\"\nFROM\n\tweb_sales T0\n\tJOIN date_dim T1\n\t\tON T0.ws_sold_date_sk = T1.d_date_sk\nGROUP BY\n\tDATE_TRUNC('month', T1.\"d_date\")::DATE", "datasource": "Warehouse", "datasource_uid": "warehouse", "adapter": "postgres" } ``` Run that SQL with your own warehouse client. 0sql has done its part. ## 9. Add a second fact table Add a `Store Sales` table with a `Net Paid` measure joined to the same `Item` and `Date` tables, deploy, and ask for both measures by month: ```sh zsql sql --expr "month(date), web net paid, net paid" ``` The SQL now has one aggregation CTE per fact table, stitched with a `FULL OUTER JOIN` on the month. Neither measure is inflated by the other's rows. That is grain-safe blending, and it needed no extra modeling. See the [cross-fact blend](#cross-fact-blend) example for the full SQL. ## Next steps - [The query spec](#the-query-spec): every key of a request. - [Request cookbook](#request-cookbook): real requests and the SQL they produce. - [Row-level security](#row-level-security): policies in the model, context in the request. - [Tests](#tests): assert the SQL a projection produces, on every deploy. ### Tutorial: TPC-DS with DuckDB This tutorial walks you through building a semantic model for the TPC-DS benchmark using **DuckDB**. You will set up the TPC-DS tutorial project, point it at a DuckDB file, deploy it to 0sql, and complete the model by adding the missing web channel tables and relationships. No cloud database is required: 0sql returns the SQL and you run it in the DuckDB shell. ## Overview By the end of this tutorial you will have: - A local DuckDB database populated with TPC-DS data - A 0sql project (cloned from [tpcds-tutorial](https://github.com/stratasite/tpcds-tutorial)) with table and relationship models for store, catalog, and inventory channels - Your own additions: the **web channel** tables and relationships (web_site, web_page, promotion, web_sales, web_returns) that are intentionally left out of the repository as a learning exercise ## Prerequisites - **Git** and the **DuckDB CLI** installed - **zsql** installed ([Installation](#installing-zsql)) and a personal API key from the console at `https://app.0sql.io` - Familiarity with [Core concepts](#core-concepts) (tables, relationships, datasources) ## Step 1: Set Up DuckDB with TPC-DS Data DuckDB is an embedded analytics database that runs as a single file. The TPC-DS benchmark is available as a DuckDB extension, so you can generate the schema and data without installing a separate database server. ### Install DuckDB Install the DuckDB CLI for your platform. See the [DuckDB installation guide](https://duckdb.org/docs/installation) for details. **Example (macOS with Homebrew):** ```bash curl https://install.duckdb.org | sh ``` ### Create the TPC-DS Database Create a new DuckDB database file and load the TPC-DS extension. Then generate the schema and data. 1. Create and open a database file (e.g. `tpcds.duckdb` in a directory of your choice): ```bash duckdb tpcds.duckdb ``` 2. In the DuckDB shell, run: ```sql INSTALL tpcds; LOAD tpcds; CALL dsdgen(sf=1); ``` This installs the TPC-DS extension, loads it, and generates data at scale factor 1. For schema only (no data), use `CALL dsdgen(sf=0);`. See the [DuckDB TPC-DS extension documentation](https://duckdb.org/docs/stable/core_extensions/tpcds) for more options. 3. Exit the DuckDB shell (e.g. `.quit` or `Ctrl+D`). You now have a file `tpcds.duckdb` containing the full TPC-DS schema and data. Note its full path; you will use it in the next step. ## Step 2: Clone the TPC-DS Tutorial Repository Clone the TPC-DS tutorial project. It already contains table and relationship models for the store, catalog, and inventory channels, plus shared dimension tables. ```bash git clone https://github.com/stratasite/tpcds-tutorial cd tpcds-tutorial ``` ### Project Structure ``` tpcds-tutorial/ ├── project.yml # Project configuration ├── datasources.yml # Datasource definitions (you will edit this) ├── models/ # Semantic model files │ ├── catalog/ # Catalog channel (catalog_sales, catalog_returns, etc.) │ ├── common/ # Shared dimensions (date_dim, customer, item, etc.) │ ├── inventory/ # Inventory fact and relationships │ └── store/ # Store channel (store_sales, store_returns, etc.) ├── security.yml # Row-level security policies └── tests/ # Planner tests, run on every deploy ``` Anything else the repository carries (for example a `migrations/` folder) is ignored by 0sql. The repository is **intentionally incomplete**: it does not include the **web channel** (web_sales, web_returns, web_site, web_page, promotion). You will add those in a later step. ## Step 3: Configure the DuckDB Datasource The cloned project may be configured for PostgreSQL. Point it at your local DuckDB file instead. 1. Open `datasources.yml`. 2. Set the `tpcds` datasource to use the DuckDB adapter and the path to your database file. Use an absolute path to avoid path issues. **Example `datasources.yml`:** ```yaml tpcds: name: TPC-DS (DuckDB) adapter: duckdb tier: hot file: /path/to/your/tpcds.duckdb ``` Replace `/path/to/your/tpcds.duckdb` with the actual path to the `tpcds.duckdb` file you created in Step 1. The service reads `name`, `adapter` and `tier`; `file` is carried as metadata for your own application, and 0sql never opens it. For more options, see the [DuckDB adapter](#duckdb-adapter) reference. Also open `project.yml`: if it has a `server:` line left over from another setup, delete it. The hosted service address is stored by `zsql auth` in the next step. ## Step 4: Sign In and Check the Project Make sure you are inside the `tpcds-tutorial` directory. Store your personal key and the service address in `.zsql` (gitignored, mode 0600): ```bash zsql auth --api-key zsk_... --server https://app.0sql.io ``` Then validate the project on the service without deploying anything: ```bash zsql check ``` The output ends with a `validated tpcds-tutorial/:` line and counts of tables, fields, joins and paths. Any problem is reported as `Error in :` followed by the message; the cloned model should pass. ## Step 5: Explore Existing Tables and Relationships Before adding new content, skim the existing models to understand the patterns. - **Fact tables** (e.g. `models/store/tbl.store_sales.yml`, `models/catalog/tbl.catalog_sales.yml`) have `cost: 100` and define measures with aggregations like `sum(ss_quantity)`. - **Dimension tables** (e.g. `models/common/tbl.date_dim.yml`, `models/common/tbl.customer.yml`) have `cost: 10` and define dimensions with `sql` expressions pointing at physical columns. - **Relationship files** (e.g. `models/store/rel.store.yml`) define `many_to_one` relationships from fact tables to dimension tables using the table display names and key columns. The same patterns apply to the web channel tables you will add next. ## Step 6: Deploy and Query Before adding your own tables, deploy the existing model and ask it for SQL. Deploys go to the checked-out git branch, so check which one you are on: ```bash git branch --show-current zsql deploy ``` The first line of output says what is being sent and where (`deploy tpcds-tutorial (... bytes) as tpcds-tutorial/ to https://app.0sql.io`), the second is the service's summary with the same counts `zsql check` printed, then one line per test. Now plan a query. The shorthand names fields, nothing else: ```bash zsql sql --expr "item category, store net paid" ``` The first output line is `-- datasource: TPC-DS (DuckDB)`; the rest is one DuckDB `SELECT` that sums `ss_net_paid` from `store_sales`, joins `item` for the category and groups by it. Paste it into `duckdb tpcds.duckdb` to run it against your file. That is the whole product: 0sql returns the SQL, your application runs it. Two more commands worth knowing before you extend the model: ```bash zsql explore --expr "store net paid" zsql explain --expr "year, store net paid, item category = Books" ``` `explore` lists every dimension and measure that can still join the query, which is how you see what the relationship files make reachable. `explain` prints the SQL plus the table route and the planner's node tree. ## Step 7: Add the Web Channel Tables The TPC-DS benchmark includes a **web channel**: fact tables `web_sales` and `web_returns`, and dimension tables `web_site`, `web_page`, and `promotion`. These are not in the tutorial repository; you will add them now. ### 7a. Add the Web Site Dimension Table Use the CLI to create the table file from the template: ```bash zsql new table "Web Site" --datasource tpcds --physical-name web_site --domain web ``` This creates `models/web/tbl.web_site.yml` with every key documented in comments. Set `cost: 10` and add dimensions for columns like `web_site_id`, `web_name`, `web_class`, etc. Follow the same style as `tbl.store.yml`. You can copy-paste the reference content below.
Complete `tbl.web_site.yml` content (click to expand) ```yaml name: Web Site physical_name: web_site datasource: tpcds cost: 10 fields: - type: dimension name: Web Site ID description: Business key for web site data_type: string expression: lookup: true sql: web_site_id - type: dimension name: Web Site Name description: Name of the web site data_type: string expression: lookup: true sql: web_name - type: dimension name: Web Site Class description: Class of the web site data_type: string expression: lookup: true sql: web_class - type: dimension name: Web Site Manager description: Manager of the web site data_type: string expression: lookup: true sql: web_manager - type: dimension name: Web Market ID description: Market identifier for the web site data_type: integer expression: lookup: true sql: web_mkt_id - type: dimension name: Web Market Class description: Market class of the web site data_type: string expression: lookup: true sql: web_mkt_class - type: dimension name: Web Market Description description: Description of the market data_type: string expression: lookup: true sql: web_mkt_desc - type: dimension name: Web Market Manager description: Manager of the market data_type: string expression: lookup: true sql: web_market_manager - type: dimension name: Web Company ID description: Company identifier for the web site data_type: integer expression: lookup: true sql: web_company_id - type: dimension name: Web Company Name description: Name of the company data_type: string expression: lookup: true sql: web_company_name - type: dimension name: Web Site City description: City where the web site is headquartered data_type: string expression: lookup: true sql: web_city - type: dimension name: Web Site County description: County where the web site is headquartered data_type: string expression: lookup: true sql: web_county - type: dimension name: Web Site State description: State where the web site is headquartered data_type: string expression: lookup: true sql: web_state - type: dimension name: Web Site ZIP description: ZIP code of the web site data_type: string expression: lookup: true sql: web_zip - type: dimension name: Web Site Country description: Country where the web site is headquartered data_type: string expression: lookup: true sql: web_country - type: dimension name: Web Site GMT Offset description: GMT offset for the web site location data_type: decimal expression: lookup: true sql: web_gmt_offset - type: dimension name: Web Site Tax Percentage description: Tax percentage applied at the web site data_type: decimal expression: lookup: true sql: web_tax_percentage ```
### 7b. Add the Web Page Dimension Table Use the CLI to create the table file: ```bash zsql new table "Web Page" --datasource tpcds --physical-name web_page --domain web ``` Edit the generated `models/web/tbl.web_page.yml` to set `cost: 10` and add dimensions for columns like `wp_web_page_id`, `wp_url`, `wp_type`, etc. You can copy-paste the reference content below.
Complete `tbl.web_page.yml` content (click to expand) ```yaml name: Web Page physical_name: web_page datasource: tpcds cost: 10 fields: - type: dimension name: Web Page ID description: Business key for web page data_type: string expression: lookup: true sql: wp_web_page_id - type: dimension name: Web Page Autogen Flag description: Whether the page was auto-generated data_type: string expression: lookup: true sql: wp_autogen_flag - type: dimension name: Web Page URL description: URL of the web page data_type: string expression: lookup: true sql: wp_url - type: dimension name: Web Page Type description: Type of web page data_type: string expression: lookup: true sql: wp_type - type: dimension name: Web Page Character Count description: Number of characters on the page data_type: integer expression: lookup: true sql: wp_char_count - type: dimension name: Web Page Link Count description: Number of links on the page data_type: integer expression: lookup: true sql: wp_link_count - type: dimension name: Web Page Image Count description: Number of images on the page data_type: integer expression: lookup: true sql: wp_image_count - type: dimension name: Web Page Max Ad Count description: Maximum number of ads on the page data_type: integer expression: lookup: true sql: wp_max_ad_count ```
### 7c. Add the Promotion Dimension Table Use the CLI to create the table file: ```bash zsql new table Promotion --datasource tpcds --domain web ``` Edit the generated `models/web/tbl.promotion.yml` to set `cost: 10` and add dimensions for columns like `p_promo_id`, `p_promo_name`, `p_channel_*`, etc. You can copy-paste the reference content below.
Complete `tbl.promotion.yml` content (click to expand) ```yaml name: Promotion physical_name: promotion datasource: tpcds cost: 10 fields: - type: dimension name: Promotion ID description: Business key for promotion data_type: string expression: lookup: true sql: p_promo_id - type: dimension name: Promotion Name description: Name of the promotion data_type: string expression: lookup: true sql: p_promo_name - type: dimension name: Promotion Cost description: Cost of running the promotion data_type: decimal expression: lookup: true sql: p_cost - type: dimension name: Promotion Response Target description: Target response rate for the promotion data_type: integer expression: lookup: true sql: p_response_target - type: dimension name: Promotion Channel Direct Mail description: Whether promotion uses direct mail channel data_type: string expression: lookup: true sql: p_channel_dmail - type: dimension name: Promotion Channel Email description: Whether promotion uses email channel data_type: string expression: lookup: true sql: p_channel_email - type: dimension name: Promotion Channel Catalog description: Whether promotion uses catalog channel data_type: string expression: lookup: true sql: p_channel_catalog - type: dimension name: Promotion Channel TV description: Whether promotion uses TV channel data_type: string expression: lookup: true sql: p_channel_tv - type: dimension name: Promotion Channel Radio description: Whether promotion uses radio channel data_type: string expression: lookup: true sql: p_channel_radio - type: dimension name: Promotion Channel Press description: Whether promotion uses press channel data_type: string expression: lookup: true sql: p_channel_press - type: dimension name: Promotion Channel Event description: Whether promotion uses event channel data_type: string expression: lookup: true sql: p_channel_event - type: dimension name: Promotion Channel Demo description: Whether promotion uses demo channel data_type: string expression: lookup: true sql: p_channel_demo - type: dimension name: Promotion Discount Active description: Whether discount is currently active data_type: string expression: lookup: true sql: p_discount_active - type: dimension name: Promotion Purpose description: Purpose of the promotion data_type: string expression: lookup: true sql: p_purpose ```
### 7d. Add the Web Sales Fact Table Use the CLI to create the fact table file: ```bash zsql new table "Web Sales" --datasource tpcds --physical-name web_sales --domain web ``` Edit the generated `models/web/tbl.web_sales.yml` to set `cost: 100` (fact tables have higher cost), and add measures with aggregations like `sum(ws_quantity)`, `sum(ws_sales_price)`, `sum(ws_net_profit)`, etc. Follow the same style as `tbl.store_sales.yml` or `tbl.catalog_sales.yml`. You can copy-paste the reference content below.
Complete `tbl.web_sales.yml` content (click to expand) ```yaml name: Web Sales physical_name: web_sales datasource: tpcds cost: 100 fields: - type: dimension name: Web Order Number description: Order number for web sale data_type: integer expression: sql: ws_order_number - type: measure name: Web Quantity description: Quantity of items sold via web data_type: integer expression: sql: sum(ws_quantity) - type: measure name: Web Wholesale Cost description: Wholesale cost for web items data_type: decimal expression: sql: sum(ws_wholesale_cost) - type: measure name: Web List Price description: List price of web items data_type: decimal expression: sql: sum(ws_list_price) - type: measure name: Web Sales Price description: Actual sales price of web items data_type: decimal expression: sql: sum(ws_sales_price) - type: measure name: Web Extended Discount Amount description: Extended discount amount for web sales data_type: decimal expression: sql: sum(ws_ext_discount_amt) - type: measure name: Web Extended Sales Price description: Extended sales price of web items data_type: decimal expression: sql: sum(ws_ext_sales_price) - type: measure name: Web Extended Wholesale Cost description: Extended wholesale cost of web items data_type: decimal expression: sql: sum(ws_ext_wholesale_cost) - type: measure name: Web Extended List Price description: Extended list price of web items data_type: decimal expression: sql: sum(ws_ext_list_price) - type: measure name: Web Extended Tax description: Extended tax amount on web sales data_type: decimal expression: sql: sum(ws_ext_tax) - type: measure name: Web Coupon Amount description: Coupon amount applied to web sales data_type: decimal expression: sql: sum(ws_coupon_amt) - type: measure name: Web Extended Ship Cost description: Extended shipping cost for web orders data_type: decimal expression: sql: sum(ws_ext_ship_cost) - type: measure name: Web Net Paid description: Net amount paid for web sales data_type: decimal expression: sql: sum(ws_net_paid) - type: measure name: Web Net Paid Including Tax description: Net paid including tax for web sales data_type: decimal expression: sql: sum(ws_net_paid_inc_tax) - type: measure name: Web Net Paid Including Ship description: Net paid including shipping for web sales data_type: decimal expression: sql: sum(ws_net_paid_inc_ship) - type: measure name: Web Net Paid Including Ship and Tax description: Net paid including shipping and tax data_type: decimal expression: sql: sum(ws_net_paid_inc_ship_tax) - type: measure name: Web Net Profit description: Net profit from web sales data_type: decimal expression: sql: sum(ws_net_profit) ```
### 7e. Add the Web Returns Fact Table Use the CLI to create the fact table file: ```bash zsql new table "Web Returns" --datasource tpcds --physical-name web_returns --domain web ``` Edit the generated `models/web/tbl.web_returns.yml` to set `cost: 100` and add measures like `sum(wr_return_quantity)`, `sum(wr_return_amt)`, `sum(wr_net_loss)`, etc. Follow the same style as `tbl.catalog_returns.yml`. You can copy-paste the reference content below.
Complete `tbl.web_returns.yml` content (click to expand) ```yaml name: Web Returns physical_name: web_returns datasource: tpcds cost: 100 fields: - type: dimension name: Web Return Order Number description: Order number for web return data_type: integer expression: sql: wr_order_number - type: measure name: Web Return Quantity description: Quantity of items returned from web data_type: integer expression: sql: sum(wr_return_quantity) - type: measure name: Web Return Amount description: Total return amount for web data_type: decimal expression: sql: sum(wr_return_amt) - type: measure name: Web Return Tax description: Tax amount on web returns data_type: decimal expression: sql: sum(wr_return_tax) - type: measure name: Web Return Amount Including Tax description: Return amount including tax for web data_type: decimal expression: sql: sum(wr_return_amt_inc_tax) - type: measure name: Web Return Fee description: Fee charged on web returns data_type: decimal expression: sql: sum(wr_fee) - type: measure name: Web Return Ship Cost description: Shipping cost for web returns data_type: decimal expression: sql: sum(wr_return_ship_cost) - type: measure name: Web Refunded Cash description: Cash refunded to customer for web returns data_type: decimal expression: sql: sum(wr_refunded_cash) - type: measure name: Web Reversed Charge description: Charge reversed on credit card for web returns data_type: decimal expression: sql: sum(wr_reversed_charge) - type: measure name: Web Account Credit description: Account credit issued for web returns data_type: decimal expression: sql: sum(wr_account_credit) - type: measure name: Web Net Loss description: Net loss from web returns data_type: decimal expression: sql: sum(wr_net_loss) ```
## Step 8: Add Relationships for the Web Channel Use the CLI to create a relationship file: ```bash zsql new relation --datasource tpcds --domain web ``` This creates `models/web/rel.web.yml`. Edit it to define `many_to_one` relationships from **Web Sales** and **Web Returns** to the dimension tables (Date, Time, Item, Customer, Web Site, Web Page, Promotion, etc.). Use the same structure as `models/store/rel.store.yml` or `models/catalog/rel.catalog.yml`: set `datasource: tpcds`, then for each relationship set `left`, `right`, `sql` (e.g. `left.ws_sold_date_sk = right.d_date_sk`), and `cardinality: many_to_one`. Use the **display names** of the tables (e.g. "Web Sales", "Date", "Item") so 0sql can resolve them. You can copy-paste the reference content below.
Complete `rel.web.yml` content (click to expand) ```yaml datasource: tpcds # Web Sales to Date Dimension (sold date) web_sales_sold_date: left: Web Sales right: Date sql: left.ws_sold_date_sk = right.d_date_sk cardinality: many_to_one # Web Sales to Time Dimension (sold time) web_sales_sold_time: left: Web Sales right: Time sql: left.ws_sold_time_sk = right.t_time_sk cardinality: many_to_one # Web Sales to Item Dimension web_sales_item: left: Web Sales right: Item sql: left.ws_item_sk = right.i_item_sk cardinality: many_to_one # Web Sales to Billed Customer web_sales_billed_customer: left: Web Sales right: Billed Customer sql: left.ws_bill_customer_sk = right.c_customer_sk cardinality: many_to_one # Web Sales to Billed Customer Demographics web_sales_billed_customer_demographics: left: Web Sales right: Billed Customer Demographics sql: left.ws_bill_cdemo_sk = right.cd_demo_sk cardinality: many_to_one # Web Sales to Billed Household Demographics web_sales_billed_household_demographics: left: Web Sales right: Billed Household Demographics sql: left.ws_bill_hdemo_sk = right.hd_demo_sk cardinality: many_to_one # Web Sales to Billed Customer Address web_sales_billed_customer_address: left: Web Sales right: Billed Customer Address sql: left.ws_bill_addr_sk = right.ca_address_sk cardinality: many_to_one # Web Sales to Shipped Customer web_sales_customer: left: Web Sales right: Customer sql: left.ws_ship_customer_sk = right.c_customer_sk cardinality: many_to_one # Web Sales to Shipped Customer Demographics web_sales_customer_demographics: left: Web Sales right: Customer Demographics sql: left.ws_ship_cdemo_sk = right.cd_demo_sk cardinality: many_to_one # Web Sales to Shipped Household Demographics web_sales_household_demographics: left: Web Sales right: Household Demographics sql: left.ws_ship_hdemo_sk = right.hd_demo_sk cardinality: many_to_one # Web Sales to Shipped Customer Address web_sales_customer_address: left: Web Sales right: Customer Address sql: left.ws_ship_addr_sk = right.ca_address_sk cardinality: many_to_one # Web Sales to Web Page web_sales_web_page: left: Web Sales right: Web Page sql: left.ws_web_page_sk = right.wp_web_page_sk cardinality: many_to_one # Web Sales to Web Site web_sales_web_site: left: Web Sales right: Web Site sql: left.ws_web_site_sk = right.web_site_sk cardinality: many_to_one # Web Sales to Promotion web_sales_promotion: left: Web Sales right: Promotion sql: left.ws_promo_sk = right.p_promo_sk cardinality: many_to_one # Web Sales to Ship Mode web_sales_ship_mode: left: Web Sales right: Ship Mode sql: left.ws_ship_mode_sk = right.sm_ship_mode_sk cardinality: many_to_one # Web Sales to Warehouse web_sales_warehouse: left: Web Sales right: Warehouse sql: left.ws_warehouse_sk = right.w_warehouse_sk cardinality: many_to_one # Web Returns to Date Dimension (returned date) web_returns_date: left: Web Returns right: Date sql: left.wr_returned_date_sk = right.d_date_sk cardinality: many_to_one # Web Returns to Time Dimension (return time) web_returns_returned_time: left: Web Returns right: Time sql: left.wr_returned_time_sk = right.t_time_sk cardinality: many_to_one # Web Returns to Item Dimension web_returns_item: left: Web Returns right: Item sql: left.wr_item_sk = right.i_item_sk cardinality: many_to_one # Web Returns to Refunded Customer web_returns_refunded_customer: left: Web Returns right: Customer sql: left.wr_refunded_customer_sk = right.c_customer_sk cardinality: many_to_one # Web Returns to Refunded Customer Demographics web_returns_refunded_customer_demographics: left: Web Returns right: Customer Demographics sql: left.wr_refunded_cdemo_sk = right.cd_demo_sk cardinality: many_to_one # Web Returns to Refunded Household Demographics web_returns_refunded_household_demographics: left: Web Returns right: Household Demographics sql: left.wr_refunded_hdemo_sk = right.hd_demo_sk cardinality: many_to_one # Web Returns to Refunded Customer Address web_returns_refunded_customer_address: left: Web Returns right: Customer Address sql: left.wr_refunded_addr_sk = right.ca_address_sk cardinality: many_to_one # Web Returns to Returning Customer web_returns_returning_customer: left: Web Returns right: Returning Customer sql: left.wr_returning_customer_sk = right.c_customer_sk cardinality: many_to_one # Web Returns to Returning Customer Demographics web_returns_returning_customer_demographics: left: Web Returns right: Returning Customer Demographics sql: left.wr_returning_cdemo_sk = right.cd_demo_sk cardinality: many_to_one # Web Returns to Returning Household Demographics web_returns_returning_household_demographics: left: Web Returns right: Returning Household Demographics sql: left.wr_returning_hdemo_sk = right.hd_demo_sk cardinality: many_to_one # Web Returns to Returning Customer Address web_returns_returning_customer_address: left: Web Returns right: Returning Customer Address sql: left.wr_returning_addr_sk = right.ca_address_sk cardinality: many_to_one # Web Returns to Web Page web_returns_web_page: left: Web Returns right: Web Page sql: left.wr_web_page_sk = right.wp_web_page_sk cardinality: many_to_one # Web Returns to Reason web_returns_reason: left: Web Returns right: Reason sql: left.wr_reason_sk = right.r_reason_sk cardinality: many_to_one ```
After creating all the files, your `models/web/` directory should contain: ``` models/web/ ├── tbl.web_site.yml ├── tbl.web_page.yml ├── tbl.promotion.yml ├── tbl.web_sales.yml ├── tbl.web_returns.yml └── rel.web.yml ``` ## Step 9: Check and Fix Issues Validate the full semantic model on the service without deploying: ```bash zsql check ``` Fix any reported errors (e.g. typos in table names, missing relationships, or invalid SQL). Keep running `zsql check` until it passes. Watch the `warning:` lines too: a table without `cost`, or two routes of equal cost, are worth fixing now. ## Step 10: Deploy and Query the Web Channel Now that you've added the web channel tables and relationships, deploy your updated model: ```bash zsql deploy ``` The summary line should show five more tables and the new joins and paths. Ask for the measures that did not exist ten minutes ago: ```bash zsql sql --expr "year, web net paid, web net profit" zsql sql --expr "item category, store net paid, web net paid" ``` The first returns a `SELECT` over `web_sales` joined to `date_dim`. The second mixes a store measure and a web measure: the planner aggregates each fact in its own universe and merges them on `item`, in one statement, because both facts reach the conformed `Item` table through the relationships you wrote. Run either in the DuckDB shell to see the numbers. Your additions have expanded the semantic model to cover all four TPC-DS channels (store, catalog, inventory, and web). See [Deploying](#deploying) for branches, `--watch` and CI. ## Summary You have: 1. Installed DuckDB and created a local TPC-DS database using the [DuckDB TPC-DS extension](https://duckdb.org/docs/stable/core_extensions/tpcds). 2. Cloned the [tpcds-tutorial](https://github.com/stratasite/tpcds-tutorial) repository and configured it to use your DuckDB file. 3. Explored the existing store, catalog, and inventory models. 4. Deployed the existing model with `zsql deploy` and read the SQL that `zsql sql` returned for it. 5. Added the missing web channel tables (web_site, web_page, promotion, web_sales, web_returns) and relationships. 6. Ran `zsql check`, deployed again, and queried the new web channel measures. ## Next Steps - [TPC-DS Tutorial (Examples)](#tpc-ds-tutorial): deeper dive into the TPC-DS model and patterns - [Query API](#query-api): POST the same requests from your application - [Star schema pattern](#star-schema-pattern): how fact and dimension tables relate - [CLI Reference](#zsql-cli): all zsql commands - [Semantic Model](#tables): tables, fields, and relationships in detail --- ## Query API The contract is small. You send a query spec (what to project, filter, calculate and segment) together with a security context describing the caller, and 0sql returns one SQL statement for the datasource the model lives in. Planning takes microseconds. Nothing is executed and nothing is stored: your application runs the SQL against its own warehouse. The same planner also runs Strata. ## Endpoint ```http POST https://app.0sql.io/projects/{uid}/branches/{branch}/sql Authorization: Bearer zqk_... Content-Type: application/json ``` - `{uid}` is the project uid from `project.yml`, `{branch}` the deployed branch. This section uses `tpcds` and `main`. - Both query keys (`zqk_`) and personal keys (`zsk_`) may call it. Query keys are read only and work on the projects and branches they are granted; personal keys do whatever their user may do. See [Accounts and keys](#accounts-keys-and-access). - A branch that has no deployment answers `404 NotFound`. ## Request envelope Send a `spec` or a one-line shorthand `expr`, plus an optional `context`. If both `spec` and `expr` are present, `spec` wins. Neither gives `400 Invalid` with the message `give a spec or an expr`. ```json {"spec": {"projections": [{"field": "category"}, {"field": "ws_net_paid"}]}, "context": {"email": "tank@matrix.com", "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}]}} ``` ```json {"expr": "category, web net paid, category = Books", "context": {"email": "tank@matrix.com"}} ``` The `context` is required on a branch that has security policies (otherwise `400 ContextRequired`) and optional elsewhere. See [Security context](#security-context). ## Response envelope ```json { "sql": "SELECT ...", "datasource": "Warehouse", "datasource_uid": "warehouse", "adapter": "postgres", "corrections": [{"term": "departmnt", "field_uid": "department", "field_name": "Department", "score": 0.8}], "spec": {"projections": [{"field": "category"}]}, "items": ["projection category"] } ``` | key | always | meaning | |---|---|---| | `sql` | yes | the single statement to run | | `datasource`, `datasource_uid`, `adapter` | yes | which warehouse the statement is for and its dialect | | `corrections` | only when non-empty | field references that were fuzzy-corrected; see [Errors and corrections](#errors-and-corrections) | | `spec` | only for `expr` requests | the spec the shorthand produced | | `items` | only for `expr` requests | how each comma-separated item was read | ## Error envelope Every error has the same shape and an HTTP status that follows the class: ```json {"error": {"class": "Semantic::NotFound", "message": "No field named 'revenue' in this model."}} ``` Spec and planner errors are `422`; missing context, bad shorthand and bad envelopes are `400`; key problems are `401` and `403`. The full table is on [Errors and corrections](#errors-and-corrections). ## A complete example Two measures from different fact tables, drilled across a date they share. `Web Net Paid` lives on `web_sales`, `Net Paid` on `store_sales`; both facts join `item`, so the category filter applies to each. **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{ "spec": { "name": "Query One", "projections": [ {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "ws_sold_date", "alias": "Web Sold Date"}, {"field": "net_paid", "alias": "Net Paid"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] } }' ``` **zsql:** ```sh zsql sql --expr 'ws_net_paid as Web Net Paid, ws_sold_date as Web Sold Date, net_paid as Net Paid, category in (men, children)' ``` The CLI parses the shorthand locally and sends a `spec`. The quoted `"books,com"` item from the JSON cannot be expressed in shorthand; see [the quoted-list gotcha](#the-quoted-list-gotcha). **JavaScript:** ```js const res = await fetch("https://app.0sql.io/projects/tpcds/branches/main/sql", { method: "POST", headers: { Authorization: `Bearer ${process.env.ZSQL_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ spec: { name: "Query One", projections: [ { field: "ws_net_paid", alias: "Web Net Paid" }, { field: "ws_sold_date", alias: "Web Sold Date" }, { field: "net_paid", alias: "Net Paid" }, ], filters: [{ field: "category", predicate: "in_list", value: 'men ,children, "books,com"' }], }, }), }); const { sql, datasource, adapter } = await res.json(); // run `sql` with your own warehouse client ``` The `sql` that comes back: ```sql WITH ag7098b0d0901f2eb16d14f9356f0bb2a0 AS ( SELECT T0."ss_sold_date_sk" AS "dimed56b67", sum(T0."ss_net_paid") AS "msr621f67c" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T0."ss_sold_date_sk" ), ag29130fae5cf6548d6ddb1bd37298b407 AS ( SELECT sum(T0."ws_net_paid") AS "msr501e4a8", T0."ws_sold_date_sk" AS "dimed56b67" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T0."ws_sold_date_sk" ) SELECT A0.msr501e4a8 AS "Web Net Paid", COALESCE(A1.dimed56b67, A0.dimed56b67) AS "Web Sold Date", A1.msr621f67c AS "Net Paid" FROM ag29130fae5cf6548d6ddb1bd37298b407 A0 FULL OUTER JOIN ag7098b0d0901f2eb16d14f9356f0bb2a0 A1 ON A1.dimed56b67 = A0.dimed56b67 ``` The two measures come from two fact tables with different grains, so summing them in one `FROM` would fan out the rows; the planner aggregates each fact in its own CTE at the shared date grain first. The final pass joins the two aggregates on that date with a `FULL OUTER JOIN` so a day with web sales but no store sales (or the reverse) still appears, with `COALESCE` picking the date from whichever side has it. Your application takes this statement and runs it against the warehouse named in `datasource`. The join type is a dialect setting (`final_pass_measure_join_type`) you can override per query through `db_settings`; see [The query spec](#the-query-spec) and [Extended blending groups](#extended-blending-groups) for how such blends are modelled. ## In this section - [The query spec](#the-query-spec): every top-level key, how field references resolve, a full annotated spec. - [Projections and decorators](#projections-and-decorators): projection keys, alias derivation, truncate, extract, temporalize, contribute, window and customize with their SQL. - [Filters](#filters): flat lists and and/or trees, every predicate, list and date values, top-n, measure filters as `HAVING`. - [Calculations](#calculations): the `[Name]@m` formula grammar, measure versus dimension calculations, complex measures. - [Segments](#segments): populations keyed on dimensions, query-level versus measure-level, exclude, expanding segments for pair analysis. - [Security context](#security-context): the context JSON and how policies turn it into `WHERE` and `CASE`. - [Shorthand expressions](#shorthand-expressions): the one-line grammar behind `expr`, the CLI and the playground. - [Explain](#explain): `POST .../explain`, timings, phases and the node tree. - [Discovery](#discovery-explore-fields-and-tables): `POST .../explore`, `GET .../fields`, `.../tables`, `.../branches/{branch}` and `GET /projects`. - [Errors and corrections](#errors-and-corrections): every class and status, the exact messages, the corrections array. Two sibling routes take the same request body: `POST .../explain` returns the SQL plus the plan ([Explain](#explain)) and `POST .../explore` returns the fields a query can still add ([Discovery](#discovery-explore-fields-and-tables)). The full route list is in the [API reference](#api-reference), and the CLI side is on [Querying with zsql](#querying-from-the-cli). ### Calculations A calculation is a column computed from a SQL formula. The formula references measures and dimensions by name in brackets and the planner substitutes the right expressions at the right node, so a ratio of two measures from two fact tables lands after both are aggregated, and a window over a count lands outside the `GROUP BY`. ## Declaring one Inline among the projections with `calculation: true`, or in the top-level `calculations` list (where every entry is a calculation and follows the projections): ```json {"projections": [ {"field": "category"}, {"field": "net_paid", "alias": "Net Paid"}, {"alias": "Bucket", "calculation": true, "sql": "[Net Paid] / 100", "data_type": "decimal"} ]} ``` ```json {"projections": [{"field": "category"}, {"field": "net_paid", "alias": "Net Paid"}], "calculations": [{"alias": "Bucket", "sql": "[Net Paid] / 100"}]} ``` | key | required | notes | |---|---|---| | `alias` | yes | non-blank: `Validation failed: Alias can't be blank` | | `sql` | yes | the formula; blank: `Sql can't be blank (calculation 'X')` | | `data_type` | no | default `decimal`; one of `string`, `integer`, `decimal`, `date`, `date_time`, `boolean`, `bigint`, `binary` | | `axis`, `order_by`, `format`, `hidden`, `decorators` | no | as on a [projection](#projections-and-decorators) | `field` and `field_type` are ignored on a calculation; its kind comes from the formula. ## Formula grammar The formula is SQL text with three kinds of bracket reference: | reference | resolves to | |---|---| | `[Name]@m` | a measure, by uid, name or synonym (fuzzy correction applies) | | `[Name]@d` | a dimension | | `[Alias]` | another projection or calculation in the same query, matched by its display alias, case-insensitive | Whitespace inside brackets is normalized: `[ Net Paid ]` is `[Net Paid]`. An `[Alias]` that matches nothing fails with: ``` Sql could not find a projection with alias X. If you meant to references a measure or dimension used @m or @d to clarify: [My Measure]@m ``` **Measure or dimension calculation.** If the formula calls an aggregate function it is a measure calculation and is computed where the query aggregates; otherwise it is a dimension calculation. The aggregates recognised: `sum`, `count`, `avg`, `min`, `max`, `median`, `mode`, `stddev`, `stddev_pop`, `stddev_samp`, `variance`, `var_pop`, `var_samp`, `array_agg`, `string_agg`, `listagg`, `bool_and`, `bool_or`, `every`, `bit_and`, `bit_or`, `count_if`, `percentile_cont`, `percentile_disc`, `corr`, `covar_pop`, `covar_samp`, `approx_distinct`, `approx_count_distinct`. **No bare columns.** Any identifier that is not a reserved keyword or a function name is taken as a raw table column and rejected: ``` Sql columns like x,y are not permitted in calculations. ``` Everything must go through brackets. String literals in single quotes are fine; double-quoted identifiers count as columns. **Allowed constructs.** `CASE WHEN ... THEN ... ELSE ... END`, `nullif`, `coalesce`, `cast`, `extract(... from ...)`, `||`, arithmetic, and window clauses `over (partition by ... order by ...)`. There is no `if()`; write `CASE`. **Calc over calc.** A calculation may reference another by `[Alias]`. A cycle fails with `Calculation X references itself through Y`. ## Examples Each of these is a real formula the planner accepts, with what it does. 1. Ratio of two measures that live on different fact tables. Each is aggregated in its own CTE and the division happens at the merge. ```json {"alias": "2x WPaid", "sql": "[Net Paid]@m/[Web Net Paid]@m", "data_type": "decimal", "calculation": true} ``` 2. Aggregated ratio: aggregates around measure references make it a measure calculation. ```json {"alias": "Agg Ratio", "sql": "sum([Web Net Paid]@m)/max([Net Paid]@m)", "calculation": true} ``` 3. Conditional distinct count: how many categories had any sale, in store or on the web. ```json {"alias": "Actives", "data_type": "integer", "calculation": true, "sql": "count(distinct case when [Net Paid]@m > 0 or [Web Net Paid]@m > 0 then [Category]@d else null end)"} ``` 4. A calculation over another calculation, with a window: each category's actives relative to the first category's. ```json {"alias": "Retention Rate", "calculation": true, "sql": "[Actives] / nullif(first_value([Actives]) over (order by [Category]@d), 0)"} ``` 5. Dimension calculation producing a string: concatenation of two dimensions. `[Category]` here resolves by alias to the projected `Category` column. ```json {"alias": "Cat Concat Item", "data_type": "string", "calculation": true, "sql": "[Category] || ' ' || [Product Name]@d"} ``` 6. Date part as a dimension calculation. ```json {"alias": "Year", "data_type": "integer", "calculation": true, "sql": "EXTRACT(year FROM [Date]@d)"} ``` 7. Margin percent, guarded against division by zero. ```json {"alias": "Margin %", "calculation": true, "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)"} ``` ## The SQL A calculation standing on a measure, inside a blend of two facts. `[Net Paid]` references the projection by alias: ```json {"spec": { "projections": [ {"field": "category", "alias": "Category"}, {"field": "net_paid", "alias": "Net Paid"}, {"alias": "Bucket", "calculation": true, "sql": "[Net Paid] / 100", "data_type": "decimal"}, {"field": "ws_net_paid", "alias": "Web Net Paid"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] }} ``` ```sql WITH agc51c5a013f468165df0d33fc12150511 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ss_net_paid") AS "msr8a51bb0" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ), ag89dc4ad5879fa4552696fe3de5b10435 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ws_net_paid") AS "msr60b3792" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ) SELECT COALESCE(A1.dim30d09b7, A0.dim30d09b7) AS "Category", A1.msr8a51bb0 AS "Net Paid", A1.msr8a51bb0 / 100 AS "Bucket", A0.msr60b3792 AS "Web Net Paid" FROM ag89dc4ad5879fa4552696fe3de5b10435 A0 FULL OUTER JOIN agc51c5a013f468165df0d33fc12150511 A1 ON A1.dim30d09b7 = A0.dim30d09b7 ``` `Bucket` is computed in the final `SELECT` over the already-aggregated `msr8a51bb0`, not inside the store sales CTE. A formula that mixed `[Net Paid]@m` and `[Web Net Paid]@m` would land in the same place, with both sides available. A calculation over a calculation, with a window. Projections: `category`, `Actives = count(distinct [Category]@d)` (integer) and `Retention Rate = [Actives] / nullif(first_value([Actives]) over (order by [Category]@d), 0)`: ```sql SELECT A0.dim30d09b7 AS "Category", count(distinct A0.dim30d09b7) AS "Actives", count(distinct A0.dim30d09b7) / nullif(first_value(count(distinct A0.dim30d09b7)) over (order by A0.dim30d09b7), 0) AS "Retention Rate" FROM ( SELECT T0."i_category" AS "dim30d09b7" FROM web_sales_sum T0 GROUP BY T0."i_category" ) A0 GROUP BY A0.dim30d09b7 ``` `[Actives]` is expanded to its formula wherever it is used, and the `first_value(...) over (...)` window stays out of the `GROUP BY`. Which table answers the category list (`web_sales_sum` in that fixture) is the resolver's choice; see [Universe formation](#universe-formation) and [Cost optimization](#cost-optimization). ## Decorators on calculations A calculation accepts `decorators` like a projection; they are validated without a field, so the kind rules apply to the calculation's own kind. A measure calculation can carry `window`, `contribute` or `temporalize`: ```json {"alias": "Margin %", "calculation": true, "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)", "decorators": [{"type": "window", "mode": "moving", "function": "avg", "size": 3}]} ``` ## Complex measures How the pieces combine. Each step is a valid formula; the SQL shape is described rather than shown. **Margin percent.** Two measures from the same fact, a guarded division. Both are aggregated in the fact's node and the division happens on the sums, so it is a weighted margin, not an average of row margins. ```json {"alias": "Margin %", "calculation": true, "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)"} ``` **Conditional distinct count.** Count members that meet a measure condition. `count(distinct ...)` makes it a measure calculation; the `CASE` reads measures and a dimension together, so the planner evaluates it where both are in scope. ```json {"alias": "Buying Customers", "data_type": "integer", "calculation": true, "sql": "count(distinct case when [Net Paid]@m > 0 then [Customer Id]@d end)"} ``` **Share of another measure across facts.** Web revenue as a fraction of store revenue. The two measures come from `web_sales` and `store_sales`; each is summed in its own aggregation CTE at the projected grain, the CTEs are joined on the shared dimensions, and the ratio is computed in the final pass (the same shape as `Bucket` above, with a measure on each side). ```json {"alias": "Web Share", "calculation": true, "sql": "[Web Net Paid]@m / nullif([Net Paid]@m + [Web Net Paid]@m, 0)"} ``` **Ratio referencing a window.** First declare the windowed quantity as its own calculation, then reference it by alias. The window stays in the final `SELECT`, outside the `GROUP BY`, like `Retention Rate` above. ```json [ {"alias": "Running Web", "calculation": true, "sql": "sum([Web Net Paid]@m) over (order by [Month(Date)])"}, {"alias": "Share of Running", "calculation": true, "sql": "[Web Net Paid]@m / nullif([Running Web], 0)"} ] ``` `[Month(Date)]` is the derived alias of a month-truncated `date` projection; give the projection an explicit alias if you prefer `[Month]`. The built-in `window` decorator covers the plain running and moving cases without a formula. ## Next steps - [Projections and decorators](#projections-and-decorators) for what a decorator can do before you reach for a formula - [Shorthand expressions](#shorthand-expressions): `Margin % := ([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)` - [Errors and corrections](#errors-and-corrections) ### Security context The security context describes the caller of a query: an email, admin flags, tags and group memberships. It travels with the request and is read by the branch's security policies to decide which rows the statement may return and which values it may show. 0sql has no users of its own; your application builds the context from its own authentication system, per request. ## The JSON ```json "context": { "email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": ["region:emea"], "groups": [ {"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]} ] } ``` | key | type | default | read by | |---|---|---|---| | `email` | string | none | policies with `source: user`, `value_from: email` | | `system_admin` | bool | `false` | bypasses policies whose `bypass.system_admin` is true (the default) | | `project_admin` | bool | `false` | bypasses policies whose `bypass.project_admin` is true | | `tags` | string[] of `key:value` | `[]` | policies with `source: user`, `value_from: tag` | | `groups` | `[{"name", "tags"}]` | `[]` | policies with `source: groups`; `name` and `tags` are each optional | Every key is optional. Unknown keys, at either level, are rejected. Nothing from the context is stored. ## When it is required A branch with at least one policy refuses to plan without a context: ```json {"error": {"class": "ContextRequired", "message": "this branch has security policies; a context is required to plan"}} ``` HTTP 400. On a branch with no policies the context is optional and the security phase is skipped. `POST .../explore` accepts a context and ignores it. An empty object `{}` is a valid context: it identifies nobody, so every fired policy resolves nothing and applies its `unresolved` rule. ## How a policy reads it A policy (authored in `security.yml`, see [Row-level security](#row-level-security)) fires on a node when the node projects a field (not a calculation) carrying one of the policy's trigger tags or listed by name. Then, per policy: 1. **Bypass.** If `bypass.system_admin` is set and the context has `system_admin: true`, or `bypass.project_admin` and `project_admin: true`, the policy is skipped. 2. **Allowed values** of the policy's context dimension are resolved from the context: | `source` | `value_from` | allowed values | |---|---|---| | `groups` | `name` | the names of the caller's groups | | `groups` | `tag` | the values of group tags whose key is the policy's `tag_key` | | `user` | `email` | the caller's email | | `user` | `tag` | the values of the caller's own `tags` with that key | Values are deduplicated. Other combinations resolve nothing. 3. **Unresolved.** Nothing resolved and the policy says `unresolved: allow`: skipped. `unresolved: deny` (the default): `filter_data` renders `WHERE 1 = 0` and `mask_data` masks every value. 4. **Apply.** `filter_data` adds `WHERE LOWER() IN ('tmnt', 'avcl')` to the node. `mask_data` wraps the triggered field: `CASE WHEN LOWER() IN (...) THEN field ELSE END`, where the mask is `mask_value` (default `#######`) for a string field and `NULL` for a numeric one. If the context dimension cannot be joined from the node that fired the policy, the planner refuses rather than return unprotected rows: ```json {"error": {"class": "Planner::SecurityPolicyError", "message": "Security policy context dimension 'Call Center ID' is not reachable in the universe for this query. Cannot safely enforce security policy."}} ``` ## Examples The branch has one policy: `filter_data`, triggered by the tag `pii`, context dimension `Call Center ID`, permissions from group tags with key `call_center_id`, `system_admin` bypass on, `unresolved: allow`. `Country` and `Employees` on `call_center` are tagged `pii`. ### Filter **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{ "spec": {"name": "security_test", "projections": [{"field": "country", "alias": "Country"}]}, "context": { "email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]}] } }' ``` **zsql:** ```sh cat > tank.json <<'EOF' {"email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]}]} EOF zsql sql --expr 'country as Country' --context tank.json ``` **JavaScript:** ```js // `user` comes from your own session or JWT const context = { email: user.email, system_admin: user.roles.includes("admin"), groups: user.callCenters.map((id) => ({ name: `CC-${id}`, tags: [`call_center_id:${id}`] })), }; const res = await fetch("https://app.0sql.io/projects/tpcds/branches/main/sql", { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify({ spec: { projections: [{ field: "country", alias: "Country" }] }, context }), }); ``` ```sql SELECT T0."cc_country" AS "Country" FROM call_center T0 WHERE LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl') ``` The two `call_center_id` tags across the caller's groups became the allowed values; the `cat:men` tag has a different key and is ignored. ### Mask The same context against a `mask_data` policy with `mask_value: "#######"`, projecting `employees` (an integer): ```json {"spec": {"projections": [{"field": "employees", "alias": "Employees"}]}, "context": {"email": "tank@matrix.com", "groups": [ {"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]}]}} ``` ```sql SELECT CASE WHEN LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl') THEN T0."cc_employees" ELSE NULL END AS "Employees" FROM call_center T0 ``` Every call center still appears as a row; the employee count is `NULL` outside the caller's centers. On a string field the `ELSE` would be `'#######'`. ### Admin bypass The first request with `"system_admin": true`: ```sql SELECT T0."cc_country" AS "Country" FROM call_center T0 ``` The policy's `bypass.system_admin` is true, so it is skipped and no `WHERE` is added. ### No context The same request with no `context` key answers 400 `ContextRequired`. ## Building the context in your application - Build it per request from your own auth system: the session's email, the roles that map to `system_admin` and `project_admin`, and the tenant, region or account memberships your policies key on as groups or tags. - API keys are not users. A query key identifies your application and what it may plan; the context identifies the person the statement is for. One key serves every end user. - Match the policy's resolution. A policy that reads `groups` by `tag` with `tag_key: call_center_id` needs groups carrying `call_center_id:` tags; names alone will not resolve, and an unresolved policy denies by default. - The comparison is lower-cased on both sides, so `TMNT` in a tag matches `tmnt` in the warehouse. - Policies apply inside segment CTEs, aggregation CTEs and decorator CTEs alike, because they fire per node. Use [Explain](#explain) to see `security_filters` and `security_masks` on each node. Authoring policies, trigger tags, `unresolved` and bypass flags are covered on [Row-level security](#row-level-security). ## Next steps - [Row-level security](#row-level-security) - [Explain](#explain) to see what a policy added - [Errors and corrections](#errors-and-corrections) ### Discovery: explore, fields and tables Five read routes describe a deployed branch without planning anything you have to run. They are what a field picker, an autocomplete or an agent calls before building a spec. Every one takes the same `Authorization: Bearer` header as `/sql` and works with query keys and personal keys. | route | returns | |---|---| | `POST .../explore` | the dimensions and measures a given query can still add | | `GET .../fields?q=&hidden=` | every field, optionally searched | | `GET .../tables` | every table with its fields | | `GET .../branches/{branch}` | the branch summary | | `GET /projects` | the projects and branches the key can see | `...` is `/projects/{uid}/branches/{branch}` throughout. ## explore Given a query, which other fields could join it without breaking it? Send the same body as `/sql` (`spec` or `expr`, optional `q`). A `context` is accepted and ignored: exploring is about the model, not about rows. **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/explore \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"expr": "category, web net paid", "q": "date"}' ``` **zsql:** ```sh zsql explore --expr 'category, web net paid' -q date ``` ``` date dim Date web_sales, store_sales, date_dim ~1.00 ws_sold_date dim Web Sold Date web_sales ~0.67 -- can add 2 dimensions, 0 measures · server 210 us ``` **JavaScript:** ```js const res = await fetch("https://app.0sql.io/projects/tpcds/branches/main/explore", { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify({ spec: currentSpec, q: userTyped }), }); const { dimensions, measures } = await res.json(); ``` ```json { "dimensions": [ {"uid": "date", "name": "Date", "data_type": "date", "description": "Calendar date", "tables": ["web_sales", "store_sales", "date_dim"], "score": 1.0}, {"uid": "ws_sold_date", "name": "Web Sold Date", "data_type": "date", "description": "", "tables": ["web_sales"], "score": 0.67} ], "measures": [], "total_us": 210 } ``` - Without `q` every addable field is listed, dimensions and measures separately, and `score` is absent. - With `q`, only fields whose name contains the term (score `1.0`) or that are trigram-near it (score at least 0.5) are kept, best first. - Hidden fields are never listed. Use it to drive a picker: after each selection, send the current spec and show only what comes back, so the user cannot build a query the planner will refuse. Agents get the same guarantee: a loop of explore, pick, explore converges on a plannable spec without guessing field names. ## fields ```sh curl -s "https://app.0sql.io/projects/tpcds/branches/main/fields?q=paid" \ -H "Authorization: Bearer zqk_..." ``` ```json {"fields": [ {"uid": "net_paid", "name": "Net Paid", "kind": "measure", "data_type": "decimal", "description": "Store sales net paid", "synonyms": ["revenue"], "tags": [], "hidden": false, "tables": ["store_sales"], "score": 1.0}, {"uid": "ws_net_paid", "name": "Web Net Paid", "kind": "measure", "data_type": "decimal", "description": "Web sales net paid", "synonyms": [], "tags": [], "hidden": false, "tables": ["web_sales"], "score": 1.0}, {"uid": "ws_paid_inc_tax", "name": "Web Paid Inc Tax", "kind": "measure", "data_type": "decimal", "description": "", "synonyms": [], "tags": [], "hidden": false, "tables": ["web_sales"], "score": 1.0} ]} ``` | param | effect | |---|---| | `q` | containing matches first (score 1), then fuzzy matches with their trigram score | | `hidden` | `true` includes hidden fields | Without `q` the whole catalogue is returned and `score` is absent. `kind` is `dimension` or `measure`. `tags` are the policy trigger tags from the model. CLI: `zsql fields paid`, `zsql fields --hidden`. ## tables ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/tables \ -H "Authorization: Bearer zqk_..." ``` ```json {"tables": [ {"uid": "store_sales", "name": "Store Sales", "physical_name": "store_sales", "cost": 100, "datasource": "warehouse", "fields": ["net_paid", "ss_sold_date", "customer_id", "..."]}, {"uid": "web_sales", "name": "Web Sales", "physical_name": "web_sales", "cost": 100, "datasource": "warehouse", "fields": ["ws_net_paid", "ws_sold_date", "..."]}, {"uid": "item", "name": "Item", "physical_name": "item", "cost": 10, "datasource": "warehouse", "fields": ["category", "product_name", "..."]}, {"uid": "date_dim", "name": "Date", "physical_name": "date_dim", "cost": 1, "datasource": "warehouse", "fields": ["date"]} ]} ``` `cost` is the model's planning cost (the resolver prefers cheaper tables); `fields` lists field uids. The table uids are what the spec's `hints` key accepts. CLI: `zsql tables`. ## branch summary ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main \ -H "Authorization: Bearer zqk_..." ``` ```json {"project": "tpcds", "branch": "main", "deployed_at": "2026-10-01 15:04", "datasources": 1, "tables": 6, "fields": 42, "joins": 5, "paths": 9, "policies": 1, "tests": 3, "warnings": []} ``` The same counts a deploy prints, minus the test results. CLI: `zsql status`. ## projects ```sh curl -s https://app.0sql.io/projects -H "Authorization: Bearer zqk_..." ``` ```json {"deployments": [ {"project": "tpcds", "branch": "main", "deployed_at": "2026-10-01 15:04", "datasources": 1, "tables": 6, "fields": 42, "joins": 5, "paths": 9, "policies": 1, "tests": 3, "warnings": []} ], "projects": [ {"uid": "tpcds", "name": "TPC-DS", "production_branch": "main", "visibility": "account", "protected_production": false, "level": "read"} ]} ``` `deployments` is one summary per deployed branch the key can read; `projects` lists each project with the caller's access `level` (`read`, `write` or `owner`). A query key sees only the projects and branches it is granted. CLI: `zsql list`. See [Accounts and keys](#accounts-keys-and-access). ## Next steps - [Shorthand expressions](#shorthand-expressions) to turn a picked field list into a request - [Explain](#explain) - [API reference](#api-reference) for the complete route list ### Errors and corrections Every failure is JSON with one shape, and the class tells you whose fault it is: the envelope, the key, the spec, or the model. Successful responses may also carry a `corrections` array when a misspelled field was silently matched. This page lists both so your application can act on them. ## The envelope ```json {"error": {"class": "Semantic::NotFound", "message": "No field named 'revenu' in this model. Did you mean: Revenue, Net Paid?"}} ``` `class` is stable and meant for code; `message` is for people and may name fields from the model. ## Classes and statuses | class | status | when | |---|---|---| | `Unauthorized` | 401 | no key, or a key that is not valid | | `Forbidden` | 403 | the key may not do this on this project or branch | | `NotFound` | 404 | no deployment for the project and branch | | `Invalid` | 400 | neither `spec` nor `expr` in the body | | `Shorthand` | 400 | the `expr` did not parse | | `ContextRequired` | 400 | the branch has policies and the request has no `context` | | `DeployError` | 400 | deploy or validate failed (not a query route) | | `Query::Spec::InvalidSpecError` | 422 | the spec's shape is wrong | | `ActiveRecord::RecordInvalid` | 422 | the spec is well-formed but a value fails validation; message prefixed `Validation failed: ` | | `Semantic::NotFound` | 422 | a field reference resolved to nothing | | `Semantic::Ambiguous` | 422 | a name is both a dimension and a measure | | `Planner::ResolutionError` | 422 | the planner cannot build the query from the model | | `Planner::SegmentDatasourceError` | 422 | a segment keyed in another datasource; an internal retry signal, normally not surfaced | | `Planner::SecurityPolicyError` | 422 | a policy's context dimension is unreachable from the query | | `Unimplemented` | 422 | a feature the planner does not plan | | `Dialect` | 500 | the base dialect failed to load | | `Accounts` | 500 | the account store failed | `400` means fix the request envelope; `401`/`403` means fix the key or its grants; `422` means fix the spec or the model; `5xx` is ours. ## Messages you will meet ### Envelope and auth | class | message | |---|---| | `Unauthorized` | `an API key is required: Authorization: Bearer ` | | `Unauthorized` | `this API key is not valid` | | `Forbidden` | `query key is not granted /` | | `Forbidden` | ` has no access to ` | | `Forbidden` | ` has access to ; this needs ` | | `Forbidden` | `query keys are read only; deploy with your personal key` | | `NotFound` | `no deployment for project branch ` | | `Invalid` | `give a spec or an expr` | | `ContextRequired` | `this branch has security policies; a context is required to plan` | ### Field references | class | message | |---|---| | `Semantic::NotFound` | `No field named 'x' in this model.` optionally followed by ` Did you mean: A, B, C?` | | `Semantic::Ambiguous` | `Field 'x' is ambiguous — more than one field answers to it. Use x@d or x@m to pick the dimension or the measure.` | ### Spec shape (`Query::Spec::InvalidSpecError`) | message | cause | |---|---| | `malformed spec: ...` | unknown key, wrong JSON type, or unparseable JSON anywhere in the spec | | `projection N needs a field` | a non-calculation projection without `field` | | `'x' is not an order_by (asc, desc)` | bad `order_by` | | `Unknown decorator type: X` | bad decorator `type` | | `unknown attribute 'x' on the decorator of ` | attribute not in the decorator vocabulary | | `'X' is not a predicate (filter on F)` | bad predicate name | | `filters must be an array or a logic tree, got ...` | bad `filters` type | | `a filter node needs a field or an and/or key, got X` | bad tree node | | `Filters here are a flat AND list — each entry must name a field; and:/or: groups are not supported in this position. ...` | an `and`/`or` group inside an array | | `Segment 'X' has no definition — ...`, `Segment 'X' needs at least one key dimension — ...`, `Unknown segment mode 'x' — use include or exclude`, `apply_to: no measure projection matches 'x' ...`, `apply_to: 'x' matches multiple projections — ...`, `join applies only to expanding segments — ...` | segment rules, listed in full on [Segments](#segments) | ### Validation (`ActiveRecord::RecordInvalid`) All prefixed `Validation failed: `. | message | cause | |---|---| | `Alias can't be blank` | calculation without an alias | | `Sql can't be blank (calculation 'X')` | calculation without a formula | | `Sql columns like x,y are not permitted in calculations.` | bare identifiers in a formula; use `[Name]@m` / `[Name]@d` | | `Sql could not find a projection with alias X. If you meant to references a measure or dimension used @m or @d to clarify: [My Measure]@m` | `[Alias]` that matches no projection | | `Axis the tooltip holds measures only` | `axis: tip` on a dimension | | `X can only be applied to dimensions` | truncate, extract or customize on a measure | | `only date/datetime types can be date truncated` | truncate or extract on a non-date dimension | | `Grain must be set`, `Extract must be set`, `Mode must be moving or running`, `Transform invalid transform type` | decorator missing its required attribute | | `Predicate predicate is not supported for .` | predicate not allowed on that field type | | `Predicate top n can only be applied to a dimension field` | `top_n` on a measure | | `Filter value can't be blank ( on )` | missing filter value | | `Filter value is not a number`, `Filter value values must be numeric` | non-numeric value on a numeric field | A bad calculation `data_type` (outside `string, integer, decimal, date, date_time, boolean, bigint, binary`) and a window `size`, `offset` or `buckets` below 1 are rejected in the same class. ### Planner | class | message | |---|---| | `Planner::ResolutionError` | `At least one projection required` | | `Planner::ResolutionError` | `No universe can resolve the query within datasource D.` The projected fields cannot be joined into one query from the model; see [Universe formation](#universe-formation). | | `Planner::ResolutionError` | `Could not find path for X` | | `Planner::ResolutionError` | `cannot parse date ...` | | `Planner::ResolutionError` | `Calculation X references itself through Y` | | `Planner::ResolutionError` | `Could not resolve segment S: ...` | | `Planner::SecurityPolicyError` | `Security policy context dimension '' is not reachable in the universe for this query. Cannot safely enforce security policy.` | | `Unimplemented` | `not implemented: custom predicate` | | `Unimplemented` | `calculations over rule-bearing measures`, `inclusion sub-plan outside the node's universe` | ### Shorthand (`Shorthand`, 400) | message | cause | |---|---| | `nothing to plan` | empty `expr` | | ``top needs a count, as in `category top 10 by revenue` `` | `top` without an integer | | `'...': unknown function f(); see .help` | unknown decorator function | ## The corrections array A field reference that matches nothing exactly is scored against every field of the wanted kind by trigram overlap of the normalized text (lower-cased, runs of non-alphanumerics collapsed to one space) against the field's name, uid and synonyms. Three constants decide what happens: | constant | value | effect | |---|---|---| | `MIN_LEN` | 4 | references shorter than four characters are never corrected | | `FLOOR` | 0.5 | the best candidate must score at least this | | `MARGIN` | 0.2 | and lead the runner-up by at least this | When both hold the candidate is used silently and the response reports it: ```json {"sql": "...", "datasource": "Warehouse", "datasource_uid": "warehouse", "adapter": "postgres", "corrections": [{"term": "departmnt", "field_uid": "department", "field_name": "Department", "score": 0.8}]} ``` | key | meaning | |---|---| | `term` | what the request said | | `field_uid`, `field_name` | what was used | | `score` | the overlap, rounded to two decimals | `corrections` is omitted when empty. When the floor or the margin is not met, the request fails with `Semantic::NotFound` and up to three candidates scoring at least 0.25 in ` Did you mean: A, B, C?`. `zsql` prints corrections to stderr before the SQL, one per line: ``` departmnt → Department (~0.8) ``` Treat a correction as a warning in an interactive tool (show the user what was substituted) and as an error in a pipeline, where a silently substituted field is a bug waiting to be noticed. ## Handling errors in your app ```js async function plan(body) { const res = await fetch(`https://app.0sql.io/projects/tpcds/branches/main/sql`, { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify(body), }); const json = await res.json(); if (!res.ok) { const { class: cls, message } = json.error; switch (cls) { case "Semantic::NotFound": case "Semantic::Ambiguous": case "Query::Spec::InvalidSpecError": case "ActiveRecord::RecordInvalid": case "Shorthand": throw new UserFixable(message); // show it; the request can be edited case "Planner::ResolutionError": case "Planner::SecurityPolicyError": case "Unimplemented": throw new ModelProblem(message); // the model or the policy needs a change case "ContextRequired": case "Invalid": case "Unauthorized": case "Forbidden": case "NotFound": throw new IntegrationBug(cls, message); // your code built the request or key wrong default: throw new Error(`${cls}: ${message}`); // 5xx: retry later, then report it } } if (json.corrections?.length) { log.warn("0sql corrected fields", json.corrections); } return json; // { sql, datasource, datasource_uid, adapter, ... } } ``` Three habits that pay off: - Branch on `class`, never on `message`. Messages name fields and may change wording. - Surface `Semantic::NotFound` with its `Did you mean` candidates and `Semantic::Ambiguous` with its `@d`/`@m` hint directly to whoever is building the query; they are written to be shown. - Log `corrections` with the request. A rename in the model can turn an exact match into a fuzzy one without any error. ## Next steps - [The query spec](#the-query-spec) - [Field references](#field-references) - [Troubleshooting](#troubleshooting) ### Explain `POST /projects/{uid}/branches/{branch}/explain` takes exactly the request `/sql` takes and returns everything `/sql` returns plus the plan: how long each phase took and the tree of nodes the statement was assembled from. Use it when the SQL is not what you expected, or to show your users where a number comes from. ## Request Same envelope, same auth, same `ContextRequired` rule: **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/explain \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"spec": {"projections": [ {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "ws_sold_date", "alias": "Web Sold Date"}, {"field": "net_paid", "alias": "Net Paid"}]}}' ``` **zsql:** ```sh zsql explain --expr 'ws_net_paid as Web Net Paid, ws_sold_date as Web Sold Date, net_paid as Net Paid' zsql sql --explain --spec query.json ``` ## Response All of the `/sql` fields (`sql`, `datasource`, `datasource_uid`, `adapter`, optional `corrections`, and `spec` + `items` for an `expr`), plus: | key | shape | |---|---| | `timings` | `{"parser_us", "plan_us", "total_us"}` | | `phases` | `[{"name", "us"}]`, in execution order | | `nodes` | `[ExplainNode]`, in post order; the last node is the root | Phase names: `resolve`, `segment_fork`, `complex_measure`, `exclusion`, `inclusion`, `snapshot`, `contribution`, `temporal`, `top_n`, `segment`, `security` (only when a context was sent), `reference`, `alias`, `sql`, `final_query`. ### Node fields | field | meaning | |---|---| | `id` | the node's number, referenced by `inputs` and `strategies` | | `alias` | the CTE name in the SQL: `ag…` for an aggregation, `seg…` for a segment population | | `kind` | `query` (a single-table plan that is the whole statement), `aggregation` (one fact aggregated at a grain, emitted as a CTE) or `merge` (joins its inputs in the final pass) | | `root` | whether this node produces the final `SELECT` | | `table` | the table the node reads, when it reads one | | `datasource` | the datasource uid | | `paths` | the join paths the resolver took to reach each projected field | | `purpose` | why the node exists, when it is not a plain projection of the spec (a contribution total, a shifted temporal side, a top-n ranking) | | `transform` | the temporal transform a shifted node implements | | `segment` | whether the node is a segment population | | `projections` | the columns the node emits | | `filters` | the filters applied at this node | | `inputs` | ids of the nodes this node consumes | | `strategies` | `[{"node", "kind"}]`: how each input is combined | | `join_type` | the join used to merge inputs (`full` by default in a blend) | | `group_by` | whether the node groups | | `security_filters` | the `WHERE` clauses policies added at this node | | `security_masks` | the fields policies wrapped in a `CASE` at this node | | `identity` | the hash the CTE alias is derived from | ## A worked response The cross-fact blend from the [Query API](#query-api) page, abbreviated: long strings are shortened and only the fields that carry information for this plan are shown. ```json { "sql": "WITH ag7098b0d0901f2eb16d14f9356f0bb2a0 AS (...), ag29130fae5cf6548d6ddb1bd37298b407 AS (...) SELECT ... FULL OUTER JOIN ...", "datasource": "Warehouse", "datasource_uid": "warehouse", "adapter": "postgres", "timings": {"parser_us": 41, "plan_us": 312, "total_us": 353}, "phases": [ {"name": "resolve", "us": 180}, {"name": "segment_fork", "us": 2}, {"name": "complex_measure", "us": 3}, {"name": "exclusion", "us": 1}, {"name": "inclusion", "us": 1}, {"name": "snapshot", "us": 1}, {"name": "contribution", "us": 1}, {"name": "temporal", "us": 1}, {"name": "top_n", "us": 1}, {"name": "segment", "us": 1}, {"name": "reference", "us": 9}, {"name": "alias", "us": 6}, {"name": "sql", "us": 88}, {"name": "final_query", "us": 17} ], "nodes": [ {"id": 1, "alias": "ag7098b0d0901f2eb16d14f9356f0bb2a0", "kind": "aggregation", "root": false, "table": "store_sales", "datasource": "warehouse", "paths": ["store_sales", "store_sales -> item"], "projections": ["Web Sold Date", "Net Paid"], "filters": ["category in_list men ,children, \"books,com\""], "inputs": [], "group_by": true, "security_filters": [], "security_masks": [], "identity": "7098b0d0..."}, {"id": 2, "alias": "ag29130fae5cf6548d6ddb1bd37298b407", "kind": "aggregation", "root": false, "table": "web_sales", "datasource": "warehouse", "paths": ["web_sales", "web_sales -> item"], "projections": ["Web Net Paid", "Web Sold Date"], "filters": ["category in_list men ,children, \"books,com\""], "inputs": [], "group_by": true, "security_filters": [], "security_masks": [], "identity": "29130fae..."}, {"id": 3, "alias": "final", "kind": "merge", "root": true, "projections": ["Web Net Paid", "Web Sold Date", "Net Paid"], "filters": [], "inputs": [2, 1], "strategies": [{"node": 2, "kind": "base"}, {"node": 1, "kind": "join"}], "join_type": "full", "group_by": false, "security_filters": [], "security_masks": []} ] } ``` Reading it: - Two `aggregation` nodes, one per fact table, each with its own `table`, `paths` and the shared category filter. That is why the SQL has two `ag…` CTEs: the measures live on different facts and must be summed separately before they can sit on one row. - The `merge` node is the `root`. Its `inputs` are the two aggregations and `join_type: "full"` is the `FULL OUTER JOIN` in the final pass (the `final_pass_measure_join_type` dialect setting, overridable through `db_settings`). - No `security` phase ran and every `security_filters` list is empty: the request carried no context and the branch has no policies. ## What to look for **Why a blend split.** More than one `aggregation` node, each with a different `table`, means the projected measures could not be served from one fact at one grain. The `projections` of each node show which measures went where. If you expected a single node, check the model's joins and blending group; see [Extended blending groups](#extended-blending-groups). **Which table was routed.** `table` and `paths` on each node show the resolver's choice among the tables that could answer the query, driven by table cost and partitions. To force a route, pass the table uid in the spec's `hints`. See [Universe formation](#universe-formation) and [Cost optimization](#cost-optimization). **What a policy added.** With a context, the `security` phase appears and each node that fired a policy lists its `security_filters` (the `WHERE LOWER(...) IN (...)` added) and `security_masks` (the fields wrapped in `CASE`). A policy that fires inside a segment or aggregation CTE shows on that node, not on the root. If a node lists a filter you did not expect, the field it projects carries a trigger tag. See [Security context](#security-context). **Where a decorator went.** Contribution totals, shifted temporal sides and top-n rankings each get their own node with a `purpose` (and a `transform` for temporal nodes), so you can see the extra CTE the decorator cost and what it groups by. ## zsql explain The CLI prints the SQL first, then the tree: ``` $ zsql explain --expr 'ws_net_paid as Web Net Paid, ws_sold_date as Web Sold Date, net_paid as Net Paid' -- datasource: Warehouse WITH ag7098b0d0901f2eb16d14f9356f0bb2a0 AS ( ... ) -- 3 nodes merge final (root) join full aggregation ag29130fae5cf6548d6ddb1bd37298b407 web_sales aggregation ag7098b0d0901f2eb16d14f9356f0bb2a0 store_sales -- server 353 us: parser 41 us, plan 312 us -- phases (us): resolve 180, segment_fork 2, ..., sql 88, final_query 17 ``` Any fuzzy corrections print first as `term → Field Name (~score)`. `--json` prints the raw response instead. In the repl, `.explain ` does the same. ## Next steps - [Discovery](#discovery-explore-fields-and-tables) for what a query can still add - [Security context](#security-context) - [Universe formation](#universe-formation) ### Filters Filters are plain values on fields. You name a field, a predicate and a value; the planner writes the `WHERE` (or `HAVING`, for a measure) with the right casing, quoting and date handling for the warehouse dialect. There is no SQL in a filter. ## Shape An array is a flat AND of leaves: ```json "filters": [ {"field": "category", "predicate": "in_list", "value": "men, children"}, {"field": "net_paid", "predicate": "greater_than", "value": "100"} ] ``` An object is a logic tree. A node with a `field` key is a leaf; otherwise its first key must be `and` or `or` holding an array of nodes, nestable to any depth: ```json "filters": {"or": [ {"field": "category", "predicate": "equals", "value": "books"}, {"and": [ {"field": "category", "predicate": "equals", "value": "music"}, {"field": "date", "predicate": "greater_than_or_equal_to", "value": "2024-01-01"} ]} ]} ``` An `and`/`or` group inside an array is rejected: ``` Filters here are a flat AND list — each entry must name a field; and:/or: groups are not supported in this position. For membership through any of several facts (e.g. bought in store OR catalog), list those measures in the segment's `measures` instead. ``` Other shape errors: `a filter node needs a field or an and/or key, got X` and `filters must be an array or a logic tree, got ...`. Filters keep their listed order. Segment filters use the same parser and the same shapes. ## Leaf keys | key | required | notes | |---|---|---| | `field` | yes | uid, name, synonym, `@d`/`@m` suffix or `{"uid": "..."}` | | `predicate` | yes | one of the wire names below, matched trimmed and lower-cased | | `value` | all but `is_null`/`is_not_null` | scalar: string, number or bool. `filter_value` is accepted as an alias. | | `value_end` | `between` | the upper bound. `filter_value_end` is accepted. | | `field_type` | no | kind hint | | `top_n_measure` | no | `top_n` only: the ranking measure | ## Predicates | wire name | meaning | allowed on | |---|---|---| | `equals`, `does_not_equal` | `=` and its negation | any field | | `is_null`, `is_not_null` | null test, no value | any field | | `in_list`, `exclude_list` | `IN (...)` and its negation over a comma-separated list | string, numeric dimension | | `contains`, `does_not_contain` | `LIKE '%x%'`, `NOT LIKE '%x%'` | string | | `starts_with`, `does_not_start_with` | `LIKE 'x%'` and its negation | string | | `ends_with`, `does_not_end_with` | `LIKE '%x'` and its negation | string | | `keyword` | `(col LIKE '%a%' OR col LIKE '%b%')` over the comma-split value | string | | `greater_than`, `greater_than_or_equal_to`, `less_than`, `less_than_or_equal_to` | `>`, `>=`, `<`, `<=` | numeric dimension, numeric measure, date, date_time | | `between` | range from `value` to `value_end` | numeric dimension, numeric measure, date, date_time | | `top_n` | ranking CTE, see below | string and numeric dimensions | | `custom` | accepted by the spec, not planned: `Unimplemented` `not implemented: custom predicate` | any | Numeric means `integer`, `decimal` or `bigint`. An unknown name fails with `'X' is not a predicate (filter on F)`. A predicate on a field type that does not allow it fails with class `ActiveRecord::RecordInvalid`: ``` Validation failed: Predicate Starts with predicate is not supported for Web Net Paid. ``` `top_n` on a measure: `Predicate top n can only be applied to a dimension field`. ## Values - `value` and `value_end` must be scalars; `null` means absent. Strings are trimmed and an empty string counts as absent. Arrays or objects fail with `a filter value must be a scalar`. - Missing where required: `Filter value can't be blank ( on )`. - On numeric fields the value must parse as a number: `Filter value is not a number`. For `in_list`/`exclude_list` every item must: `Filter value values must be numeric`. `top_n` needs a numeric N. ### Strings String comparisons lower-case both sides, so `Books`, `books` and `BOOKS` match the same rows: ```json {"spec": {"projections": [{"field": "ws_net_paid"}], "filters": [{"field": "category", "predicate": "starts_with", "value": "super"}]}} ``` ```sql SELECT sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") LIKE 'super%' ``` `equals` renders `LOWER(T1."i_category") = 'books'`. ### Lists A list is one string, split on commas. Each item is trimmed and lower-cased. Commas inside single or double quotes do not split, and a quoted item keeps its quotes in the SQL: ```json {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""} ``` ```sql LOWER(T1."i_category") IN ('men', 'children', '"books,com"') ``` The gotcha: quoting keeps `books,com` together, but the quotes become part of the compared value, so this matches a category literally stored as `"books,com"`. Values that contain commas and no quotes in the data cannot be expressed in a list; use `equals` leaves under an `or` tree instead. ### Numbers ```json {"field": "net_paid", "predicate": "between", "value": "100", "value_end": "500"} {"field": "employees", "predicate": "in_list", "value": "10, 20, 30"} ``` Numbers may be sent as JSON numbers or as strings. ### Null checks No `value`: ```json {"field": "date", "predicate": "is_not_null"} ``` ### Dates Fixed dates in any of these formats: `%Y-%m-%d %H:%M:%S`, `%Y-%m-%dT%H:%M:%S`, `%Y-%m-%d %H:%M`, `%Y/%m/%d %H:%M:%S`, `%Y-%m-%d`, `%Y/%m/%d`, `%d-%m-%Y`, `%b %d %Y`, `%B %d %Y`, `%d %b %Y`. Relative dates are one to three digits followed by a unit: `h`, `d`, `w`, `m`, `q`, `y`, meaning that many units ago. `w`, `m`, `q` and `y` snap to the start of the unit, or to its end when used as `value_end` of a `between`; `h` and `d` do not snap. ```json {"field": "date", "predicate": "greater_than_or_equal_to", "value": "28d"} {"field": "date", "predicate": "between", "value": "3m", "value_end": "1m"} {"field": "date", "predicate": "between", "value": "1y", "value_end": "1d"} ``` `3m` to `1m` is from the first day of the month three months ago to the last day of last month. Anything the parser cannot read is a `Planner::ResolutionError` with `cannot parse date ...`. ### top_n Keep the N dimension values that rank highest. `value` is N; the ranking measure goes in `top_n_measure` (uid or name) or is packed into the value as `"N:measure"`. With no measure the ranking is by row count. ```json {"field": "category", "predicate": "top_n", "value": "5", "top_n_measure": "ws_net_paid"} {"field": "category", "predicate": "top_n", "value": "5:ws_net_paid"} ``` Shorthand: `category top 5 by web net paid`. The planner ranks in a CTE and inner-joins the winners back: ```json {"spec": {"projections": [{"field": "ws_net_paid"}, {"field": "category"}], "filters": [ {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}, {"field": "category", "predicate": "top_n", "value": "5", "top_n_measure": "ws_net_paid"} ]}} ``` ```sql WITH ag5c9bb8d550197b3a219ae50f0b427905 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ws_net_paid") AS "msr0c0d023" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ORDER BY sum(T0."ws_net_paid") desc LIMIT 5 ) SELECT sum(T0."ws_net_paid") AS "Web Net Paid", T1."i_category" AS "Category" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN ag5c9bb8d550197b3a219ae50f0b427905 AS A2 ON T1."i_category" = A2.dim30d09b7 WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ``` The other filters apply inside the ranking CTE too, so the top five are the top five among the listed categories. This `LIMIT 5` is the only `LIMIT` the planner emits; the spec's `limit` key is not written into the statement. ### keyword `keyword` is a multi-term contains: the value is comma-split and each term becomes a `LIKE '%term%'`, joined with `OR`. ```json {"field": "product_name", "predicate": "keyword", "value": "super, ultra"} ``` ## Measure filters There is no separate syntax. A filter whose field is a measure is a measure filter, and the planner places it per node: where a node groups, dimension filters go to `WHERE` and measure filters to `HAVING`. On a grouped merge node the measure filter is wrapped in `max(...)`. ```json {"spec": {"projections": [ {"field": "date", "order_by": "asc", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "ws_net_paid", "decorators": [{"type": "temporalize", "transform": "month_over_month"}]} ], "filters": [ {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}, {"field": "ws_net_paid", "predicate": "greater_than", "value": "100"} ]}} ``` ```sql WITH ag47a9128558e109f9cf0b93c01992e0e7 AS ( SELECT (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) AS "dim5b35e57", sum(T0."ws_net_paid") AS "__msrlm_51898eff8" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) HAVING sum(T0."ws_net_paid") > 100 ) SELECT A0.dim5b35e57 AS "Month(Date)", A0.__msrlm_51898eff8 AS "LM(Web Net Paid)" FROM ag47a9128558e109f9cf0b93c01992e0e7 A0 ORDER BY A0.dim5b35e57 asc ``` The category filter is a `WHERE` on the rows; the measure filter is a `HAVING` on the month's sum. To filter a population by a measure and then report other measures over it, use a [segment](#segments) instead; a measure filter inside a segment becomes a `HAVING` on the segment CTE. ## Validation messages | message | cause | |---|---| | `'X' is not a predicate (filter on F)` | unknown predicate name | | `Validation failed: Predicate predicate is not supported for .` | predicate not allowed on that field type | | `Predicate top n can only be applied to a dimension field` | `top_n` on a measure | | `Filter value can't be blank ( on )` | missing `value` (or `value_end` for `between`) | | `Filter value is not a number` | non-numeric value on a numeric field | | `Filter value values must be numeric` | non-numeric list item on a numeric field | | `a filter value must be a scalar` | array or object as a value | | `cannot parse date ...` | unreadable date literal (`Planner::ResolutionError`) | | `not implemented: custom predicate` | `custom` predicate (`Unimplemented`) | ## Next steps - [Segments](#segments) for population membership, including OR across facts - [Shorthand expressions](#shorthand-expressions) for the one-line filter syntax - [Errors and corrections](#errors-and-corrections) ### Projections and decorators A projection is one output column: a field, optionally decorated. Decorators change how the column is computed (truncate a date to months, turn a measure into a running sum or a percent of total) without you writing SQL. This page covers the projection keys and every decorator. ## ProjectionSpec ```json {"field": "ws_net_paid", "alias": "Web Net Paid", "order_by": "desc", "decorators": [{"type": "window", "mode": "running", "function": "sum"}]} ``` | key | type | default | notes | |---|---|---|---| | `field` | string or `{"uid": "..."}` | required | a uid, a name or a synonym, optionally suffixed `@d` / `@m`. Missing: `projection N needs a field`. See [field references](#field-references). | | `field_type` | string | none | kind hint, same vocabulary as the `@` suffix | | `alias` | string | derived | the output column name; trimmed, blank means none | | `order_by` | `asc` or `desc` | none | case-insensitive. Anything else: `'x' is not an order_by (asc, desc)`. There is no top-level sort key. | | `axis` | `row`, `x`, `y`, `series`, `tip`, `pivot` | `row` | carried for your client's layout, not planned. `tip` on a dimension fails: `Axis the tooltip holds measures only`. | | `hidden` | bool | `false` | honoured: the projection takes part in the query but is not a visible column | | `format` | any JSON | none | carried back untouched | | `decorators` | `DecoratorSpec[]` | `[]` | see below | | `calculation`, `sql`, `data_type` | | | make the entry a calculation; see [Calculations](#calculations) | Projections keep their listed order. Entries in the separate top-level `calculations` list follow them. A date projection with no `truncate` decorator is the raw column; nothing is truncated for you. ## Alias derivation An explicit `alias` wins. Otherwise the highest-priority decorator lends a label; otherwise the field's name is used. | decorator | lent alias | |---|---| | truncate | `Month(Date)` (capitalized grain); `raw` lends the field name | | extract | `mon(Date)`, with the part abbreviated: minute `min`, hour `hr`, day_of_month `dom`, day_of_week `dow`, day_of_year `doy`, week_of_year `woy`, month `mon`, quarter `qtr`, year `yr`, day_name `day`, month_name `month`, year_month `ym` | | contribute | `% Web Net Paid of Total` (`Web Net Paid of Total` if the name already holds `%`) | | temporalize | `LY(Web Net Paid)`, `LQ(...)`, `LM(...)`, `LW(...)`, `D/D(...)`; prefixed `%` when `percent_change` | | window | `Running Sum(Web Net Paid)`, `Moving Avg(Web Net Paid)` | | customize | `Custom(Category)` | Decorator priorities: truncate 0, extract 1, contribute 2, temporalize 3, window 3, customize 10000. The highest wins the alias. Calculations reference other columns by this display alias (`[Month(Date)]`), so give an explicit alias when the derived one is awkward. ## Decorators A decorator is `{"type": "", ...attributes}`. Kinds: `truncate`, `extract`, `temporalize`, `contribute`, `window`, `customize` (alias `custom`). Anything else: `Unknown decorator type: X`. Only the attributes listed under each kind may be written; any other key fails with `unknown attribute 'x' on the decorator of `. Kind rules: - `truncate`, `extract`, `customize` apply to dimensions only: `X can only be applied to dimensions`. `truncate` and `extract` also need a date or datetime field: `only date/datetime types can be date truncated`. - `window`, `contribute`, `temporalize` apply to measures only. Two decorators of the same type on one projection are merged: the attributes of the second extend the first. Different types stack; `yoy(month(date))` from the shorthand produces a `truncate` and a `temporalize` on the same projection (and then fails lowering, since temporalize is measure-only). All SQL below is the engine's output for the TPC-DS example model with the filter `category in_list "men ,children, \"books,com\""` present, which is why every statement carries the same `WHERE`. ### truncate Bucket a date dimension to a grain. ```json {"type": "truncate", "grain": "month"} ``` Shorthand: `month(date)`. Grains: `raw`, `millisecond`, `second`, `minute`, `hour`, `day`, `week`, `month`, `quarter`, `year`. Required: `Grain must be set`. ```json {"spec": {"projections": [ {"field": "ws_net_paid"}, {"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}]}} ``` ```sql SELECT sum(T0."ws_net_paid") AS "Web Net Paid", DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY DATE_TRUNC('month', T1."d_date")::DATE ``` ### extract Pull one part out of a date dimension. ```json {"type": "extract", "extract": "day_of_week"} ``` Shorthand: `dow(date)`, `month_name(date)`. Parts: `minute`, `hour`, `day_of_month`, `day_of_week`, `day_of_year`, `week_of_year`, `month`, `quarter`, `year`, `day_name`, `month_name`, `year_month`. Required: `Extract must be set`. The column is grouped like any dimension; the derived alias is the abbreviation from the table above (`dow(Date)`). ### temporalize Compare a measure to the same measure one period earlier. ```json {"type": "temporalize", "transform": "month_over_month"} {"type": "temporalize", "transform": "year_over_year", "percent_change": true} ``` Shorthand: `mom(web net paid)`, `yoy_pct(web net paid)`. Transforms: `year_over_year`, `quarter_over_quarter`, `month_over_month`, `week_over_week`, `day_over_day`. Required: `Transform invalid transform type`. `percent_change` (bool) returns the relative change instead of the prior-period value. The planner builds the prior period by shifting the date in its own CTE, then joins it back at the projected grain. Month over month, with an ordered month projection and a measure filter: ```json {"spec": {"projections": [ {"field": "date", "order_by": "asc", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "ws_net_paid", "decorators": [{"type": "temporalize", "transform": "month_over_month"}]} ], "filters": [ {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}, {"field": "ws_net_paid", "predicate": "greater_than", "value": "100"} ]}} ``` ```sql WITH ag47a9128558e109f9cf0b93c01992e0e7 AS ( SELECT (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) AS "dim5b35e57", sum(T0."ws_net_paid") AS "__msrlm_51898eff8" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) HAVING sum(T0."ws_net_paid") > 100 ) SELECT A0.dim5b35e57 AS "Month(Date)", A0.__msrlm_51898eff8 AS "LM(Web Net Paid)" FROM ag47a9128558e109f9cf0b93c01992e0e7 A0 ORDER BY A0.dim5b35e57 asc ``` Each month's `LM(Web Net Paid)` is the previous month's sum: the CTE adds `'1 month'::interval` to every sold date so that January's rows land under February. Because the base measure is not itself projected, the statement is only the shifted side. Project `ws_net_paid` as well and the plan gains a second CTE for the current period, joined on the month. ### contribute A measure as a share of its total. ```json {"type": "contribute"} {"type": "contribute", "partition_refs": ["category"]} ``` Shorthand: `pct(web net paid)`, `share(web net paid, category)`. Attributes: `partition_refs` (projected dimension uids that bound the total, or `-1` for auto; refs that are not projected are dropped) and `ignore_partition_filters` (default `true`). With no partition the total is the whole measure: ```json {"spec": {"projections": [ {"field": "ws_net_paid", "decorators": [{"type": "contribute"}]}, {"field": "category"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}]}} ``` ```sql WITH ag61495e144c0531efb642c4545df2e6c5 AS ( SELECT sum(T0."ws_net_paid") AS "msr9f2f6d5", T1."i_category" AS "dim841154e" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ), agbe9574d2ae0d53a3a40c210cf88ff8c4 AS ( SELECT sum(T0."ws_net_paid") AS "__msrtotal_51898eff8" FROM web_sales T0 ) SELECT A0.msr9f2f6d5 / (A1.__msrtotal_51898eff8*1.00) AS "% Web Net Paid of Total", A0.dim841154e AS "Category" FROM ag61495e144c0531efb642c4545df2e6c5 A0 CROSS JOIN agbe9574d2ae0d53a3a40c210cf88ff8c4 A1 ``` The total CTE has no `WHERE`: `ignore_partition_filters` is true, so the share is of all web sales, not of the three filtered categories. The `*1.00` forces decimal division. ### window Running and moving aggregates, ranks, lag and lead. ```json {"type": "window", "mode": "running", "function": "sum"} {"type": "window", "mode": "moving", "function": "avg", "size": 7} {"type": "window", "mode": "running", "function": "lag", "offset": 2, "partition_refs": ["date"], "order_by": {"category": "desc"}} ``` Shorthand: `running_sum(web net paid)`, `moving_avg(web net paid, 7)`, `running_sum(web net paid, category)`. | attribute | default | rule | |---|---|---| | `mode` | required | `moving` or `running`: `Mode must be moving or running` | | `function` | required | `avg`, `sum`, `min`, `max`, `count`, `row_number`, `rank`, `dense_rank`, `percent_rank`, `cume_dist`, `ntile`, `lag`, `lead`, `first_value`, `last_value` | | `size` | 7 | `moving` mode frame, in rows; must be at least 1 | | `offset` | 0 | `lag` / `lead`; must be at least 1 | | `buckets` | 4 | `ntile`; must be at least 1 | | `partition_refs` | none | projected dimension uids for `PARTITION BY`, or `-1` for auto; non-projected refs are dropped | | `order_by` | `{"-1": "asc"}` | map of projected dimension uid to `asc`/`desc`; `-1` means the first projected date. Keys that are not projected dimensions are dropped. | | `ignore_partition_filters` | `true` | carried | A moving average of 14 renders as `avg(sum(x)) OVER (ORDER BY ROWS BETWEEN 14 PRECEDING AND CURRENT ROW)`. The lag example, with month and category projected: ```json {"spec": {"projections": [ {"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "category"}, {"field": "ws_net_paid", "decorators": [ {"type": "window", "mode": "running", "function": "lag", "offset": 2, "partition_refs": ["date"], "order_by": {"category": "desc"}}]} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}]}} ``` ```sql SELECT DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", T2."i_category" AS "Category", lag(sum(T0."ws_net_paid"), 2) OVER (PARTITION BY DATE_TRUNC('month', T1."d_date")::DATE ORDER BY T2."i_category" desc) AS "Running Lag(Web Net Paid)" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY DATE_TRUNC('month', T1."d_date")::DATE, T2."i_category" ``` The window sits over the aggregate in the same `SELECT`: `partition_refs: ["date"]` becomes `PARTITION BY` the truncated date, the `order_by` map becomes `ORDER BY ... desc`. ### customize Wrap a dimension's expression in your own SQL. ```json {"type": "customize", "custom_sql": "upper(@expression)"} ``` `custom_sql` is required, must contain `@expression` (replaced by the dimension's column expression) and must not contain `avg(`, `sum(`, `min(`, `max(` or `count(`. The derived alias is `Custom(Category)`. Dimensions only. There is no shorthand function for it. ## Attribute reference Every decorator starts from one attribute row; a kind reads the attributes it needs and the rest keep their defaults. | attribute | default | read by | |---|---|---| | `grain` | null | truncate | | `extract` | null | extract | | `mode`, `function`, `size`, `offset`, `buckets` | null, null, 7, 0, 4 | window | | `order_by` | `{"-1": "asc"}` for window | window | | `partition_refs` | null | window, contribute | | `ignore_partition_filters` | true | window, contribute | | `transform`, `percent_change` | null | temporalize | | `abbreviate` | true | carried | | `custom_sql` | null | customize | Keys named `id`, `projection_id`, `prompt`, `share_full_result_set_with_ai`, `type`, `created_at`, `updated_at`, `created_by_id`, `updated_by_id` are accepted and ignored. ## Next steps - [Filters](#filters) - [Calculations](#calculations) for anything a decorator cannot express - [Shorthand expressions](#shorthand-expressions) for the function names behind each decorator ### Segments A segment is a population: the set of values of one or more key dimensions whose members satisfy some filters, found through some measures' fact tables. The planner builds it as a CTE grouped on the keys and joins it into the query. Use segments for "customers who bought books", "products sold in store", "the top 50 stores by revenue", and for putting a cohort beside the baseline in one result. ## SegmentSpec ```json "segments": [{ "name": "Book Products", "mode": "include", "keys": ["product_name"], "measures": [], "filters": [{"field": "category", "predicate": "equals", "value": "books"}], "apply_to": ["ws_net_paid"] }] ``` | key | type | default | notes | |---|---|---|---| | `name` | string | `Segment N` (1-based) | used in messages | | `keys` | field refs | required | the dimensions identifying a member: the join grain. Must be dimensions. | | `measures` | field refs | `[]` | measures whose fact tables define membership. Must be measures. | | `filters` | array or tree | none | same shapes as the [query filters](#filters) | | `mode` | `include` or `exclude` | `include` | `exclude` is an anti-join. Not for expanding segments. | | `apply_to` | string[] | `[]` | measure projections the segment constrains; empty constrains the whole query | | `expanding` | `[{"field", "as"}]` | `[]` | dimensions of the other members sharing the key, carried into the query | | `join` | `inner` or `left` | `inner` | expanding segments only | | `pair_dedupe` | `neq`, `lt`, `none` | `neq` | expanding segments only | Each key and measure is a field reference (uid, name, synonym, `@d`/`@m`). An ambiguous key is retried as a dimension and an ambiguous measure as a measure. Unknown keys inside a segment are rejected as `malformed spec: ...`. ### Rules and messages | rule | message | |---|---| | A segment needs a definition: at least one of `keys`, `measures`, `filters`, `expanding` | `Segment 'X' has no definition — a stateless spec cannot adopt a stored segment by name; give its keys, measures and filters.` | | At least one key | `Segment 'X' needs at least one key dimension — list the member-identifying dimension UIDs under keys (e.g. keys: ["customer-id"]).` | | Keys are dimensions | `Segment key 'x' is a measure — list it under measures instead` | | `mode` is `include` or `exclude` | `Unknown segment mode 'x' — use include or exclude` | | `apply_to` names a measure projection | `apply_to: no measure projection matches 'x' (measure projections: A, B). Reference one of those verbatim, or set an alias on the projection to constrain.` | | `apply_to` is unambiguous | `apply_to: 'x' matches multiple projections — set an alias on the one to constrain and use it` | | `join` only with `expanding` | `join applies only to expanding segments — segment 'X' has no expanding dimensions` | | `mode` and `apply_to` are not allowed on an expanding segment | rejected as an invalid spec | The first message explains the design: there are no stored segments. A spec carries the full definition every time. ## Query-level segments With `apply_to` empty the segment constrains every row of the query. Web revenue by product, for products in the books category: ```json {"spec": { "name": "Product Sales", "projections": [{"field": "product_name", "alias": "Product Name"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}], "segments": [{"name": "Book Products", "mode": "include", "keys": ["product_name"], "filters": [{"field": "category", "predicate": "equals", "value": "books"}]}] }} ``` ```sql WITH seg4a9429dd0d8da260126c0bdb19cc5c6a AS ( SELECT T0."i_product_name" AS "dimbe52306" FROM item T0 WHERE LOWER(T0."i_category") = 'books' GROUP BY T0."i_product_name" ) SELECT T1."i_product_name" AS "Product Name", sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN seg4a9429dd0d8da260126c0bdb19cc5c6a AS Q2 ON T1."i_product_name" = Q2.dimbe52306 GROUP BY T1."i_product_name" ``` The `seg…` CTE is the population: one row per key value that passes the filter. The query `INNER JOIN`s it on the key. Here the segment only needs `item`, because the filter and the key both live there; a segment whose filter needs a fact table (or that lists `measures`) is built from that fact. ## Measure-level segments `apply_to` names the measure projections the segment constrains, by display alias first, then by measure uid or name. Other measures stay unconstrained, which puts the cohort beside the baseline in one result. Same query plus store revenue, with the books segment applied only to the web measure: ```json {"spec": { "name": "Product Sales", "projections": [ {"field": "product_name", "alias": "Product Name"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "net_paid", "alias": "Net Paid"} ], "segments": [{"name": "Book Products", "mode": "include", "keys": ["product_name"], "filters": [{"field": "category", "predicate": "equals", "value": "books"}], "apply_to": ["ws_net_paid"]}] }} ``` ```sql WITH ag6610e03993b3da57101097f16f2df44d AS ( SELECT T1."i_product_name" AS "dimbe52306", sum(T0."ss_net_paid") AS "msr621f67c" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_product_name" ), seg4a9429dd0d8da260126c0bdb19cc5c6a AS ( SELECT T0."i_product_name" AS "dimbe52306" FROM item T0 WHERE LOWER(T0."i_category") = 'books' GROUP BY T0."i_product_name" ), ag4e173d8139dc2fe7bf4cf4d6aced1d58 AS ( SELECT T1."i_product_name" AS "dimbe52306", sum(T0."ws_net_paid") AS "msr0c0d023" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN seg4a9429dd0d8da260126c0bdb19cc5c6a AS Q2 ON T1."i_product_name" = Q2.dimbe52306 GROUP BY T1."i_product_name" ) SELECT COALESCE(A1.dimbe52306, A0.dimbe52306) AS "Product Name", A0.msr0c0d023 AS "Web Net Paid", A1.msr621f67c AS "Net Paid" FROM ag4e173d8139dc2fe7bf4cf4d6aced1d58 A0 FULL OUTER JOIN ag6610e03993b3da57101097f16f2df44d A1 ON A1.dimbe52306 = A0.dimbe52306 ``` The segment join sits only inside the web sales aggregation CTE. The store sales CTE has no segment, so `Net Paid` covers every product, and the `FULL OUTER JOIN` keeps products that appear on either side. One request gives the cohort's measure and the baseline measure side by side per product. To compare the same measure in and out of the cohort, project it twice with different aliases and point `apply_to` at one of them; `apply_to` matches the alias verbatim. ## include and exclude `mode: "exclude"` keeps the rows whose key is not in the population (an anti-join). Products never sold in the books category: ```json {"name": "Not Books", "mode": "exclude", "keys": ["product_name"], "filters": [{"field": "category", "predicate": "equals", "value": "books"}]} ``` ## Measures in a segment Listing a measure under `measures` routes membership through that measure's fact table, and a filter on a measure becomes a `HAVING` on the segment CTE. Products with more than 1000 in store revenue: ```json {"name": "Store Sellers", "keys": ["product_name"], "measures": ["net_paid"], "filters": [{"field": "net_paid", "predicate": "greater_than", "value": "1000"}]} ``` The CTE groups `store_sales` by product and applies `HAVING sum(ss_net_paid) > 1000`; the query then inner-joins the surviving products. A `top_n` filter inside a segment ranks the population the same way: the top 50 products by store revenue as a segment, then any measure over them. This is also how to express membership through any of several facts. A flat filter list cannot say "bought in store or on the web" (the error for an `or` group inside a flat list says as much); list both measures instead: ```json {"name": "Any Channel", "keys": ["customer_id"], "measures": ["net_paid", "ws_net_paid"]} ``` ## Expanding segments An expanding segment carries a dimension of the other members that share the key into the query, which turns a fact into a pair matrix: products bought by the same customer, items in the same order. `expanding` lists the carried dimensions; `as` (wire name) is the output column, defaulting to `" (same )"`. ```json {"spec": { "projections": [{"field": "product_name"}, {"field": "ws_net_paid"}], "segments": [{"name": "Customer Products", "keys": ["customer_id"], "expanding": [{"field": "product_name", "as": "Product Name (same Customer ID)"}]}] }} ``` ```sql WITH seg87ae4658940c714d42c8e2714d387839 AS ( SELECT T1."c_customer_id" AS "dimf5dcb79", T2."i_product_name" AS "dimd3303bd" FROM store_sales T0 JOIN customer T1 ON T0.ss_customer_sk = T1.c_customer_sk JOIN item T2 ON T0.ss_item_sk = T2.i_item_sk GROUP BY T1."c_customer_id", T2."i_product_name" ) SELECT T1."i_product_name" AS "Product Name", sum(T0."ws_net_paid") AS "Web Net Paid", Q2.dimd3303bd AS "Product Name (same Customer ID)" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN customer T3 ON T0.ws_bill_customer_sk = T3.c_customer_sk INNER JOIN seg87ae4658940c714d42c8e2714d387839 AS Q2 ON T3."c_customer_id" = Q2.dimf5dcb79 AND T1."i_product_name" <> Q2.dimd3303bd GROUP BY T1."i_product_name", Q2.dimd3303bd ``` The CTE is every (customer, product) pair. The query joins it on the customer and keeps the companion product as a column, so each row is "web revenue of product A, for customers who also bought product B". The `<>` is the default `pair_dedupe: "neq"`, which drops self-pairs. | option | values | effect | |---|---|---| | `join` | `inner` (default), `left` | `left` keeps rows whose key has no companion, with a null carried column | | `pair_dedupe` | `neq` (default), `lt`, `none` | `neq` drops A with A; `lt` keeps one canonical ordering of each pair; `none` keeps everything | `mode` and `apply_to` cannot be combined with `expanding`. ## Next steps - [Filters](#filters) for the predicates a segment filter can use - [Security context](#security-context): policies apply inside segment CTEs too - [Example requests](#request-cookbook) ### Shorthand expressions The shorthand is a one-line form of the spec: comma-separated items, each a projection, a filter or a calculation. It exists for people typing queries (the CLI, the repl, the playground) and for agents that find a line easier to produce than JSON. Everything it can say, the JSON spec can say; the reverse is not true. ``` month(date) as Month asc, category, web net paid, category in (Books, Music), date >= 1y ``` ## Where it is accepted - `{"expr": "..."}` in place of `spec` in `POST .../sql`, `.../explain` and `.../explore`. The response then carries `spec` (the generated spec) and `items` (one line per item saying how it was read). Parse errors come back as `400` class `Shorthand`. - `zsql sql --expr '...'`, `zsql explain --expr`, `zsql explore --expr`. The CLI parses locally and sends a `spec`, printing the item descriptions to stderr. `zsql spec '...'` prints the spec without calling the server. - The repl (`zsql` inside a project): any line not starting with `{` is shorthand. - The playground in the console. ## Grammar **Splitting.** The line is split on commas that are outside quotes, square brackets and parentheses. An empty line is `nothing to plan`. **Classification** of each item, in order: 1. **Escaped field.** `[Name]`, `[Name]@m`, `[Name]@d` (also `@dim`, `@measure`), `"Name"` or `'Name'` as the whole item is a projection of that field and nothing else. The inner text may not contain `[`, `]`, `"` or `'`. 2. **Calculation.** `Alias = formula` or `Alias := formula`, where the formula contains `[`, the alias is non-empty, contains none of `< > ! =`, and is not itself a bracketed or quoted field. Produces `{"alias", "sql", "calculation": true}`. Operator and keyword searches never look inside bracketed or quoted spans. 3. **Filter.** Checked on the lower-cased text, in this order: ` is not null`, ` is null` (suffixes); ` not in (...)`, ` in (...)` (items split on commas, each unquoted, re-joined with `, `); ` between a and b`; ` top N [by measure]` (N must be an integer, else ``top needs a count, as in `category top 10 by revenue` ``); ` contains `, ` starts with `, ` ends with `, each with an optional ` not ` or ` does not ` before the keyword; then the operators `>=`, `<=`, `!=`, `<>`, `=`, `>`, `<`. The value is unquoted. 4. **Projection.** Optional trailing ` asc` / ` desc` becomes `order_by`; optional ` as Alias` becomes `alias`; then nested decorator functions `f(g(field, args...))`, emitted innermost first in `decorators`. An unknown function fails with `'...': unknown function f(); see .help`. 5. **Unnamed calculation.** If that fails and the text contains `[`, it is a calculation: `formula as Alias`, or, with no alias, the formula is its own alias. Filter predicate names emitted: `is_not_null`, `is_null`, `exclude_list`, `in_list`, `between`, `top_n`, `contains`, `does_not_contain`, `starts_with`, `does_not_start_with`, `ends_with`, `does_not_end_with`, `greater_than_or_equal_to`, `less_than_or_equal_to`, `does_not_equal`, `equals`, `greater_than`, `less_than`. There is no `and`/`or` in the shorthand. Filters are always a flat AND list; `web net paid > 100 and category = Books` parses as one filter on a field called `web net paid > 100 and category`. Use the JSON tree for OR. ## Decorator functions | shorthand | JSON | |---|---| | `raw(x)`, `millisecond(x)`, `second(x)`, `minute(x)`, `hour(x)`, `day(x)`, `week(x)`, `month(x)`, `quarter(x)`, `year(x)` | `{"type": "truncate", "grain": ""}` | | `dow(x)`, `dom(x)`, `doy(x)`, `woy(x)`, `hr(x)`, `min(x)`, `mon(x)`, `qtr(x)`, `yr(x)`, `day_name(x)`, `month_name(x)`, `year_month(x)`, and the long forms `day_of_week(x)`, `day_of_month(x)`, `day_of_year(x)`, `week_of_year(x)` | `{"type": "extract", "extract": ""}` | | `yoy(x)`, `qoq(x)`, `mom(x)`, `wow(x)`, `dod(x)` | `{"type": "temporalize", "transform": "year_over_year" ...}` | | `yoy_pct(x)`, `qoq_pct(x)`, `mom_pct(x)`, `wow_pct(x)`, `dod_pct(x)` | the same with `"percent_change": true` | | `pct(x)`, `share(x)`, `contribute(x)`, `contribution(x)`, each optionally `(x, dim, ...)` | `{"type": "contribute"}` plus `"partition_refs": [dims]` when given | | `running_(x[, dims...])` | `{"type": "window", "mode": "running", "function": ""}` plus `"partition_refs"` for the dims | | `moving_(x, N[, dims...])` | `{"type": "window", "mode": "moving", "function": "", "size": N}` plus `"partition_refs"` | `` is any window function: `sum`, `avg`, `min`, `max`, `count`, `row_number`, `rank`, `dense_rank`, `percent_rank`, `cume_dist`, `ntile`, `lag`, `lead`, `first_value`, `last_value`. There is no shorthand for the `customize` decorator. The parser does not type-check. `yoy(month(date))` parses fine and then fails lowering, because `temporalize` is measure-only. Errors of that kind come back as `422` with the spec's validation message, not as class `Shorthand`. ## Examples Each line with the spec it produces, as emitted by the parser. 1. Two projections. ``` category, web net paid ``` ```json {"projections": [{"field": "category"}, {"field": "web net paid"}]} ``` 2. A truncated date, an alias and a sort. ``` month(date), ws_net_paid as Web Paid desc ``` ```json {"projections": [ {"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "ws_net_paid", "alias": "Web Paid", "order_by": "desc"}]} ``` 3. Nested functions, innermost first in the output. This one parses but fails lowering (temporalize on a dimension). ``` yoy(month(date)) as Growth desc ``` ```json {"projections": [{"field": "date", "alias": "Growth", "order_by": "desc", "decorators": [{"type": "truncate", "grain": "month"}, {"type": "temporalize", "transform": "year_over_year"}]}]} ``` 4. Percent change and an equality filter. ``` yoy_pct(web net paid), category = Books ``` ```json {"projections": [{"field": "web net paid", "decorators": [{"type": "temporalize", "transform": "year_over_year", "percent_change": true}]}], "filters": [{"field": "category", "predicate": "equals", "value": "Books"}]} ``` 5. A list and a date range. Filters alone are a valid parse; the server then fails with `At least one projection required`. ``` category in (Books, Music), date between 2024-01-01 and 2024-12-31 ``` ```json {"projections": [], "filters": [ {"field": "category", "predicate": "in_list", "value": "Books, Music"}, {"field": "date", "predicate": "between", "value": "2024-01-01", "value_end": "2024-12-31"}]} ``` 6. A quoted value and a null check. ``` product name contains 'super', date is not null ``` ```json {"projections": [], "filters": [ {"field": "product name", "predicate": "contains", "value": "super"}, {"field": "date", "predicate": "is_not_null"}]} ``` 7. Top N by a measure. ``` category top 5 by web net paid ``` ```json {"projections": [], "filters": [{"field": "category", "predicate": "top_n", "value": "5", "top_n_measure": "web net paid"}]} ``` 8. A named calculation. ``` Share = sum([Web Net Paid]@m) / sum([Net Paid]@m) ``` ```json {"projections": [], "calculations": [{"alias": "Share", "sql": "sum([Web Net Paid]@m) / sum([Net Paid]@m)", "calculation": true}]} ``` 9. A calculation named with `as`. ``` max([net paid]@m) as Peak ``` ```json {"projections": [], "calculations": [{"alias": "Peak", "sql": "max([net paid]@m)", "calculation": true}]} ``` 10. Contribution, running sum, moving average. ``` pct(web net paid), running_sum(web net paid), moving_avg(web net paid, 7) ``` ```json {"projections": [ {"field": "web net paid", "decorators": [{"type": "contribute"}]}, {"field": "web net paid", "decorators": [{"type": "window", "mode": "running", "function": "sum"}]}, {"field": "web net paid", "decorators": [{"type": "window", "mode": "moving", "function": "avg", "size": 7}]}]} ``` 11. Escaped fields: brackets for a name that would otherwise be read as something else, quotes for a field whose name contains an operator word. ``` [Tables], "Sales in Store" > 5 ``` ```json {"projections": [{"field": "Tables"}], "filters": [{"field": "Sales in Store", "predicate": "greater_than", "value": "5"}]} ``` 12. Negations. ``` category != Books, net paid >= 100, name not starts with x ``` ```json {"projections": [], "filters": [ {"field": "category", "predicate": "does_not_equal", "value": "Books"}, {"field": "net paid", "predicate": "greater_than_or_equal_to", "value": "100"}, {"field": "name", "predicate": "does_not_start_with", "value": "x"}]} ``` 13. `:=` for a calculation whose alias contains a symbol. ``` Margin % := ([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0) ``` ```json {"projections": [], "calculations": [{"alias": "Margin %", "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)", "calculation": true}]} ``` 14. Extract parts. ``` dom(date), month_name(date) as Mon ``` ```json {"projections": [ {"field": "date", "decorators": [{"type": "extract", "extract": "day_of_month"}]}, {"field": "date", "alias": "Mon", "decorators": [{"type": "extract", "extract": "month_name"}]}]} ``` ## The quoted-list gotcha The shorthand strips the quotes of list items, so a quoted item that contains a comma is later split by the planner: ``` category in(Books,"Music, Live") ``` ```json {"field": "category", "predicate": "in_list", "value": "Books, Music, Live"} ``` That is three values, not two. For list items containing commas, send a JSON spec and quote the item inside the value string (see [lists](#lists)). ## Sending an expr **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"expr": "month(date), category, web net paid, category = Books"}' ``` **zsql:** ```sh zsql sql --expr 'month(date), category, web net paid, category = Books' ``` **JavaScript:** ```js const res = await fetch("https://app.0sql.io/projects/tpcds/branches/main/sql", { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify({ expr: "month(date), category, web net paid, category = Books" }), }); const { sql, spec, items } = await res.json(); ``` The response includes the generated `spec` and `items` such as `projection date` and `filter category equals Books`, so a client can show the user how the line was read before running the SQL. ## Next steps - [The query spec](#the-query-spec) for everything the shorthand cannot express: `or` trees, segments, `db_settings`, `customize` - [Querying with zsql](#querying-from-the-cli) for the repl and its dot-commands ### The query spec A spec is a JSON object describing one query: which fields to project, how to decorate them, what to filter, which calculations and segments to add. It is stateless; the server knows nothing about it before the request and nothing after. This page is the top-level reference. Each sub-object has its own page. ## Top-level keys | key | type | required | default and notes | |---|---|---|---| | `name` | string | no | carried onto the query, not planned | | `description` | string | no | carried | | `limit` | integer | no | 5000 when absent. The value is carried, but the returned statement has no `LIMIT` clause; the only `LIMIT` you will see belongs to a `top_n` ranking CTE. Apply your own limit when you run the SQL. | | `projections` | [`ProjectionSpec[]`](/docs/query/projections/) | yes in effect | an empty list fails with `Planner::ResolutionError` `At least one projection required` | | `calculations` | [`ProjectionSpec[]`](/docs/query/calculations/) | no | each entry is treated as `calculation: true` and ordered after the projections | | `filters` | array or object | no | a flat AND array, or an `and`/`or` tree. See [Filters](#filters). | | `segments` | [`SegmentSpec[]`](/docs/query/segments/) | no | populations joined into the query | | `view` | any JSON | no | visualization settings for your client; carried, never planned | | `db_settings` | object | no | per-query dialect overrides. Any dialect key can be set, for example `final_pass_measure_join_type` (the join between aggregation CTEs in a blend, `full` by default). The planner also reads the boolean `force_group_by`. | | `hints` | string[] | no | table uids the resolver must route through. See [Universe formation](#universe-formation). | Unknown top-level keys are rejected with class `Query::Spec::InvalidSpecError` and a message starting `malformed spec: `. The same applies to unknown keys inside a projection, a filter leaf or a segment. There is no top-level sort key: ordering is per projection through `order_by`. ```json {"spec": {"projections": [{"field": "category"}], "sort": "category"}} ``` ```json {"error": {"class": "Query::Spec::InvalidSpecError", "message": "malformed spec: unknown field `sort`, expected one of ..."}} ``` ## Field references Wherever a spec names a field (`field` in a projection or filter, segment `keys` and `measures`, `top_n_measure`, `[Name]@m` inside a formula) the reference is a string, or `{"uid": "..."}`. Resolution runs in this order: 1. Trim, then split once on `@`. The suffix (`@d`, `@dim`, `@dimension`, `@m`, `@measure`) or an explicit `field_type` key picks the kind. Anything whose last segment starts with `d` is a dimension; any other non-empty value is a measure. No suffix means either kind. 2. Exact **uid** match within the allowed kinds. 3. Else **name** match, trimmed, case-insensitive. 4. Else **synonym** match, trimmed, case-insensitive. 5. Zero matches: `Semantic::NotFound` with `No field named 'x' in this model.` More than one: `Semantic::Ambiguous`. Hidden fields resolve like any other. These three projections all point at the same measure: ```json [{"field": "ws_net_paid"}, {"field": "web net paid"}, {"field": "Web Net Paid@m"}] ``` ### Ambiguity A model may hold a dimension and a measure with the same name (`Net Paid` as the raw column and `Net Paid` as its sum). A bare reference then fails: ```json {"error": {"class": "Semantic::Ambiguous", "message": "Field 'Net Paid' is ambiguous — more than one field answers to it. Use Net Paid@d or Net Paid@m to pick the dimension or the measure."}} ``` Fix it with the suffix or with `field_type`: ```json {"field": "Net Paid@m"} {"field": "Net Paid", "field_type": "measure"} ``` ### Fuzzy correction When a reference of four or more characters matches nothing, every field of the wanted kind is scored by trigram overlap against its name, uid and synonyms. If the best score is at least 0.5 and leads the runner-up by at least 0.2, it is taken silently and reported in the response's `corrections` array: ```json {"spec": {"projections": [{"field": "categry"}, {"field": "web net paid"}]}} ``` ```json {"sql": "...", "datasource": "...", "datasource_uid": "...", "adapter": "...", "corrections": [{"term": "categry", "field_uid": "category", "field_name": "Category", "score": 0.83}]} ``` Otherwise the `Semantic::NotFound` message gains ` Did you mean: A, B, C?` with up to three candidates. References shorter than four characters never correct; they resolve exactly or fail. Details on [Errors and corrections](#errors-and-corrections). ## An annotated spec Monthly web revenue by category over the last year, as a share of the total and as month-over-month change, restricted to products that sold in store, with a margin calculation and the top five categories only. ```json { "name": "Monthly web revenue, top categories", "limit": 500, "projections": [ {"field": "date", "alias": "Month", "order_by": "asc", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "category"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "ws_net_paid", "alias": "Share of Total", "decorators": [{"type": "contribute", "partition_refs": ["date"]}]}, {"field": "ws_net_paid", "decorators": [{"type": "temporalize", "transform": "month_over_month", "percent_change": true}]}, {"alias": "Margin %", "calculation": true, "data_type": "decimal", "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)"} ], "filters": [ {"field": "date", "predicate": "between", "value": "1y", "value_end": "1d"}, {"field": "category", "predicate": "top_n", "value": "5", "top_n_measure": "ws_net_paid"} ], "segments": [ {"name": "Sold in store", "keys": ["product_name"], "measures": ["net_paid"], "filters": [{"field": "net_paid", "predicate": "greater_than", "value": "0"}]} ], "db_settings": {"final_pass_measure_join_type": "left"}, "view": {"client_state": "carried back untouched"} } ``` Line by line: - `date` with a `truncate` decorator becomes `DATE_TRUNC('month', ...)`; without the decorator a date projection is the raw column. The explicit alias `Month` overrides the derived `Month(Date)`. - The three `ws_net_paid` projections are the same measure three times: plain, as a percent of the month's total (`partition_refs` restricts the total to each month), and as a month-over-month percent change. The last one carries no `alias`, so its decorator lends one: `LM(Web Net Paid)`, prefixed with `%` because `percent_change` is set. See [Projections and decorators](#projections-and-decorators). - `Margin %` is an inline calculation: plain SQL around `[Name]@m` measure references, computed over the aggregated measures. See [Calculations](#calculations). - The `between` filter uses relative dates: `1y` snaps to the start of last year, `1d` is yesterday. `top_n` keeps the five categories ranked by `ws_net_paid`. See [Filters](#filters). - The segment is a population of `product_name` values with positive store sales; listing `net_paid` under `measures` routes membership through `store_sales`. With no `apply_to` it constrains the whole query. See [Segments](#segments). - `db_settings` turns the blend's final join into a `LEFT JOIN`; `view` and `limit` are carried back to your client untouched. The SQL for this spec is a blend (web and store measures) with a segment CTE, a contribution total CTE, a shifted-date CTE for the month-over-month measure and a top-n ranking CTE, merged in one final `SELECT`. Each piece is shown on its own page with the engine's actual output. ## Sending it Wrap the spec in the request envelope and POST it; the shape of the HTTP exchange is on the [Query API](#query-api) page. **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d @spec.json ``` where `spec.json` holds `{"spec": {...}, "context": {...}}`. **zsql:** ```sh zsql sql --spec spec.json --context ctx.json ``` `--spec` accepts the bare spec or a file with a top-level `spec` key. `zsql spec ''` prints the spec a shorthand line produces without calling the server. **JavaScript:** ```js const res = await fetch("https://app.0sql.io/projects/tpcds/branches/main/sql", { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify({ spec, context }), }); if (!res.ok) { const { error } = await res.json(); throw new Error(`${error.class}: ${error.message}`); } const { sql } = await res.json(); ``` ## Next steps - [Projections and decorators](#projections-and-decorators) - [Filters](#filters) - [Shorthand expressions](#shorthand-expressions) for the one-line form of the same spec - [Example requests](#request-cookbook) --- ## CLI Reference `zsql` is the command line for 0sql. It scaffolds a project (YAML files describing your warehouse tables, dimensions, measures, joins and security policies), validates and deploys it to the hosted service, and sends queries to it. The CLI generates no SQL itself: every plan runs on the service, and `zsql` prints what comes back. ## Install ```sh curl -fsSL https://0sql.io/install.sh | sh zsql --version ``` A single static binary, no runtime. Details and the build-from-source route are in [Installing zsql](#installing-zsql). ## The loop ```sh zsql init tpcds # project.yml, datasources.yml, security.yml, models/, tests/ cd tpcds $EDITOR datasources.yml # one entry per warehouse: adapter, tier, host... zsql new table "Store Sales" --datasource tpcds --physical-name store_sales --domain store zsql new relation --datasource tpcds --domain store zsql auth --api-key zsk_... --server https://app.0sql.io zsql check # validate on the service, keep nothing zsql deploy # deploy the checked-out git branch zsql sql --expr "item category, store net paid" --context user.json zsql repl --context user.json zsql test ``` `init` creates the files, `new` adds model files from templates, `auth` stores your personal key in `.zsql`, `check` validates without deploying, `deploy` ships the checked-out git branch, and `sql`, `repl` and `test` exercise the deployment. Each step has a page below. ## Commands | Group | Command | Does | |---|---|---| | project | `zsql init [DIR]` | Start a project: templates, `.gitignore`, `git init -b main` | | | `zsql new table NAME --datasource DS` | Write `models//tbl..yml` from the template | | | `zsql new relation --datasource DS` | Write `models//rel..yml` | | | `zsql new test NAME` | Write `tests/.yml` | | service | `zsql auth --api-key KEY` | Save the key (and optionally the server) in `.zsql` | | | `zsql check` | Validate the project on the service without deploying | | | `zsql deploy` | Deploy the checked-out branch | | | `zsql test` | Run the deployed branch's tests on the service | | | `zsql status` | The deployed branch's summary, as JSON | | | `zsql list` | Every deployment the key can see | | | `zsql remove --yes` | Remove the deployed branch | | | `zsql health` | Liveness of the server | | querying | `zsql sql` | Plan one query and print the SQL | | | `zsql explain` | SQL plus the node graph and timings | | | `zsql explore` | What a query can add: dimensions to drill into, measures to join | | | `zsql spec LINE` | The spec JSON a shorthand line makes, locally | | | `zsql repl` | Plan interactively | | | `zsql fields [Q]` | Fields of the deployed branch | | | `zsql tables` | Tables of the deployed branch | Every flag is in the [command reference](#command-reference). ## Global flags | Flag | Default | |---|---| | `--project DIR` | The directory holding `project.yml`, here or in a parent | | `--branch NAME` | The checked-out git branch of the project; `main` outside a project | | `--server URL` | `ZSQL_SERVER`, else `server:` in `.zsql` or `~/.zsql/config`, else `server:` in `project.yml` | | `--uid UID` | `uid` from `project.yml` | The API key comes from `ZSQL_API_KEY`, else `api_key:` in `.zsql` (written by `zsql auth`), else `~/.zsql/config`. Outside a project directory, `--uid` is required and the branch defaults to `main`; that is how you query a deployment from a machine that has no checkout. Without any server setting, `zsql` falls back to a local development address (`http://127.0.0.1:3699`), so pass `--server https://app.0sql.io` to `zsql auth` once. See [Authentication and configuration](#authentication-and-configuration). Server errors print as `: `, for example `NotFound: no deployment for project tpcds branch main`, and the process exits non-zero. ## Pages - [Projects](#projects): `zsql init`, `project.yml`, the directory layout, branches - [Datasources](#datasources): `datasources.yml`, adapters, tiers, secrets - [Modeling with zsql](#modeling-with-zsql): `zsql new table`, `zsql new relation`, `zsql check` - [Authentication and configuration](#authentication-and-configuration): keys, `.zsql`, environment variables, server resolution - [Deploying](#deploying): `zsql deploy`, `--watch`, `status`, `list`, `remove`, branch semantics - [Tests](#tests): `tests/*.yml` and `zsql test` - [Querying from the CLI](#querying-from-the-cli): `sql`, `explain`, `explore`, `spec`, `fields`, `tables`, the repl - [CI/CD](#cicd): GitHub Actions and GitLab CI - [Command reference](#command-reference): every command, flag, env var and file ### Authentication and configuration `zsql` authenticates every request with `Authorization: Bearer `. The key comes from an environment variable or a small YAML file in the project; the server URL comes from a flag, an environment variable, the same files or `project.yml`. ## Keys Keys are created in the console at `https://app.0sql.io`. | Kind | Prefix | For | Can | |---|---|---|---| | Personal key | `zsk_` | People and CI | Everything its user may do: deploy, test, remove, query | | Query key | `zqk_` | Applications | Read only: `sql`, `explain`, `explore`, `fields`, `tables` and the branch summary, on the projects and branches it is granted | Each key's secret is shown once, 44 characters long, and listed afterwards by its first 12 characters. `zsql` works with either kind, but `zsql deploy` with a query key answers `Forbidden: query keys are read only; deploy with your personal key`. Members, roles, grants and rotation are on [Accounts](#accounts-keys-and-access). ## zsql auth ```sh zsql auth --api-key KEY [--server URL] ``` Writes the key (and the server, if given) into `.zsql` in the project directory, replacing any existing line for the same key. The file is created with mode `0600` and is in the `.gitignore` that `zsql init` wrote. ```sh $ zsql auth --api-key zsk_1kJ9... --server https://app.0sql.io saved to /home/you/tpcds/.zsql ``` The file: ```yaml # zsql local configuration: api key and server. Do not commit. api_key: zsk_1kJ9... server: https://app.0sql.io ``` `zsql auth` must run inside a project (or with `--project DIR`), because the file is per project. One project can hold one key at a time; switch keys by running it again. ## Environment variables `ZSQL_API_KEY` wins over every file. It is the right place for a key in CI and in any shell where you do not want a file on disk: ```sh export ZSQL_API_KEY=zsk_1kJ9... zsql deploy --branch main ``` `ZSQL_SERVER` sets the server the same way. ## ~/.zsql/config The same YAML keys as `.zsql` (`api_key`, `server`), read when the project file has no value for them. Put a personal key there once and every project on the machine uses it unless its own `.zsql` says otherwise. The repl keeps its history beside it, in `~/.zsql/history`. ## Resolution order The API key: 1. `ZSQL_API_KEY` 2. `api_key:` in `.zsql` in the project directory, or `.strata` if `.zsql` does not exist 3. `api_key:` in `~/.zsql/config` The server: 1. `--server URL` 2. `ZSQL_SERVER` 3. `server:` in `.zsql` (or `.strata`) in the project directory, then in `~/.zsql/config` 4. `server:` in `project.yml` 5. `http://127.0.0.1:3699`, a local development default Because of step 5, set the server once: `zsql auth --api-key KEY --server https://app.0sql.io`, or `export ZSQL_SERVER=https://app.0sql.io`. A trailing `/` is trimmed. `.strata` is read only when `.zsql` is absent. It exists so a project that already has a `.strata` file with an `api_key` works without a second file; new projects use `.zsql`. ## zsql health ```sh zsql health ``` Resolves the server the same way and calls its unauthenticated liveness route. No key needed. ```sh $ zsql health https://app.0sql.io ok ``` A wrong or missing key shows up on the first authenticated command instead: ```sh $ zsql status Unauthorized: an API key is required: Authorization: Bearer ``` ## Next steps - [Accounts](#accounts-keys-and-access): members, keys, grants - [Deploying](#deploying) - [CI/CD](#cicd) ### CI/CD A project is a git repository, so the pipeline is short: install `zsql`, put a personal key in `ZSQL_API_KEY`, run `zsql check` on pull requests and `zsql deploy` on merge. The deployed branch is whatever `--branch` names, and the job fails when a test fails because `zsql` exits non-zero. ## What you need - A **personal key** (`zsk_...`) for a user with write access to the project, stored as a CI secret. Query keys cannot deploy. Create keys in the console at `https://app.0sql.io` (see [Accounts](#accounts-keys-and-access)); a dedicated "ci" key is easy to revoke. - `ZSQL_API_KEY` set from that secret. It wins over any `.zsql` file, and nothing is written to disk. - `ZSQL_SERVER=https://app.0sql.io`, unless `project.yml` or a committed config names the server. - **`--branch`.** CI checkouts are usually detached (`git rev-parse --abbrev-ref HEAD` says `HEAD`), and `zsql` takes the branch from git. Pass the CI's branch variable explicitly, always. ## GitHub Actions `.github/workflows/zsql.yml`: ```yaml name: zsql on: pull_request: push: branches: [main] env: ZSQL_API_KEY: ${{ secrets.ZSQL_API_KEY }} ZSQL_SERVER: https://app.0sql.io jobs: check: if: github.event_name == 'pull_request' runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Install zsql run: curl -fsSL https://0sql.io/install.sh | sh - name: Validate and run tests run: zsql check --branch "${{ github.head_ref }}" deploy: if: github.event_name == 'push' runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Install zsql run: curl -fsSL https://0sql.io/install.sh | sh - name: Deploy run: zsql deploy --branch "${{ github.ref_name }}" ``` On a pull request, `zsql check` validates the model and runs the tests on the service without deploying; the branch name only labels the output. On a push to `main`, `zsql deploy --branch main` updates `tpcds/main`. To also deploy feature branches as isolated previews, add their pattern to `push.branches` and the same step deploys `tpcds/`; remove them with `zsql remove --yes --branch ` when the pull request closes. ## GitLab CI `.gitlab-ci.yml`: ```yaml variables: ZSQL_SERVER: https://app.0sql.io default: image: debian:bookworm-slim before_script: - apt-get update -qq && apt-get install -y -qq curl ca-certificates git - curl -fsSL https://0sql.io/install.sh | sh check: stage: test rules: - if: $CI_PIPELINE_SOURCE == "merge_request_event" script: - zsql check --branch "$CI_MERGE_REQUEST_SOURCE_BRANCH_NAME" deploy: stage: deploy rules: - if: $CI_COMMIT_BRANCH == "main" script: - zsql deploy --branch "$CI_COMMIT_BRANCH" ``` Store `ZSQL_API_KEY` as a masked, protected CI/CD variable; it reaches the job as an environment variable, which is where `zsql` looks first. ## Failing on tests `zsql check` and `zsql deploy` exit non-zero when the archive is rejected or when any test fails. In the deploy case the model is already live (the HTTP call returned 200 with `test_results.failed > 0`), so a red job after a merge means "deployed, and a test broke": fix forward. To keep a broken model off `main`, require the `check` job on pull requests before merging. ``` deploy tpcds (6240 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies FAILED catalog quantity by month --- generated ... tests: 2 passed, 1 failed Error: tests failed ``` ## Protected production branches A project owner can mark the production branch protected in the console. Then only an owner's key may deploy to it, and anyone else gets: ``` Forbidden: deploying to the protected production branch 'main' of tpcds needs an owner ``` Give the CI key's user owner access on the project, or keep production deploys to a job that uses an owner's key. Everyone with write access can still deploy feature branches. ## Next steps - [Deploying](#deploying) - [Tests](#tests) - [Authentication and configuration](#authentication-and-configuration) ### Datasources `datasources.yml` describes each warehouse a project's tables live in. The `adapter` picks the SQL dialect the service emits; the `tier` tells the planner which copy of the data to prefer. 0sql never connects to a warehouse: the connection fields ride along as metadata for your own application, and secrets never leave your machine. ## The file The commented template `zsql init` writes: ```yaml # The warehouses this project's tables live in. One entry per datasource, # keyed by a short id that table files refer to. Connection details that # are secrets (password, tokens, keys) are stripped before deploy, and # nothing here is ever used by 0sql. # # warehouse: # name: Warehouse # adapter: postgres # duckdb | postgres | snowflake | bigquery | databricks | trino | mysql | sqlserver | ... # tier: hot # hot | warm | cold; the planner prefers hotter tiers # host: localhost # port: 5432 # database: analytics # username: analyst # schema: public # # local: # name: Local DuckDB # adapter: duckdb # tier: hot # file: ./data/analytics.duckdb ``` A filled-in file for the running TPC-DS example: ```yaml tpcds: name: TPC-DS (Postgres) adapter: postgres tier: cold host: localhost username: tpcds_user database: tpcds ``` The top level maps a **key** to its settings. The key (trimmed, lower-cased) is the datasource uid; table files refer to it with `datasource: tpcds`. The file is required, needs at least one entry, and refuses a repeated key. | Key | Required | Meaning | |---|---|---| | `name` | no | Display name. Defaults to the key. Printed by `zsql sql` as `-- datasource: `. | | `adapter` | yes | One of the identifiers below. Selects the dialect. | | `tier` | no | `hot`, `warm` or `cold`. The planner prefers hotter tiers when several datasources can answer a query. | | `description` | no | Free text. | | `host`, `port`, `database`, `username`, `schema`, `file` | no | Connection metadata for your own application. Not validated, and never used by 0sql. | ## Adapters ``` athena bigquery clickhouse databricks druid duckdb mysql postgres redshift snowflake sqlite sqlserver trino ``` Any other string fails the deploy with `Datasource errors: adapter '' is not supported`. The adapter controls the SQL the service emits for that datasource: identifier quoting (`"` for most, backticks for `bigquery`, `clickhouse`, `databricks` and `mysql`, `[...]` for `sqlserver`), date literals and truncation functions, which words are reserved and which functions count as aggregates in your expressions, and join support (`mysql` has no `FULL JOIN`; `druid` has no CTEs). There are no adapter-specific settings to write in the project; the adapter name is the whole configuration. ## Tiers `tier` ranks copies of the same data. When a measure is reachable on a `hot` datasource and a `cold` one, the planner routes to the hot one. Use it when a fast store holds a subset or an aggregate of what the warehouse holds. Routing rules are on [Semantic routing](#semantic-routing). ## Secrets The service never opens a connection, so it never needs credentials. On `zsql deploy`, `datasources.yml` is shipped with these keys removed: ``` password private_key personal_access_token secret_access_key access_key_id oauth_client_secret oauth_client_id api_key ``` Keep secrets out of the file anyway. Nothing in the project needs a working credential, because 0sql never opens a connection; the fields are there for your application to read. ## Next steps - [Datasources in the semantic model](#datasources) - [Modeling with zsql](#modeling-with-zsql) - [Semantic routing](#semantic-routing) ### Deploying `zsql deploy` packs the project directory into a tar.gz, POSTs it to the service under the checked-out git branch, and prints what the service loaded: counts, warnings and test results. `zsql check` does the same against the validate route and keeps nothing. ## zsql check ```sh zsql check ``` Identical to `zsql deploy --dry-run`: the same archive, the same loader, the same output with `validate`/`validated` in place of `deploy`/`deployed`. Nothing is stored. See [Modeling with zsql](#modeling-with-zsql) for a transcript. ## zsql deploy ```sh zsql deploy [--dry-run] [--watch] ``` ```sh $ git branch --show-current main $ zsql deploy deploy tpcds (6214 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies PASSED store net paid by category PASSED catalog quantity by month PASSED date range tests: 3 passed, 0 failed ``` The first line says what is being sent and where: the project name, the archive size, the `uid/branch` it lands on, and `(the production branch)` when the branch equals `production_branch` in `project.yml`. The second line is the service's summary. Then one `warning:` line per warning, one line per test, and the test total. A project with no `tests/*.yml` prints `no tests deployed (tests/*.yml)` instead. ### What goes in the archive | Shipped | Never shipped | |---|---| | `project.yml` | `.zsql`, `.strata` | | `security.yml` | `.git` | | `datasources.yml` with secret keys removed (`password`, `private_key`, `personal_access_token`, `secret_access_key`, `access_key_id`, `oauth_client_secret`, `oauth_client_id`, `api_key`) | data files, anything not listed on the left | | every `.yml` / `.yaml` under `models/` and `tests/` | | ### Exit codes `zsql deploy` exits non-zero when the service rejects the archive (a `DeployError` naming the file and message) and when any test fails. Note the second case: the HTTP deploy succeeds with status 200 and the branch is live with the new model, but the CLI prints `tests failed` and exits non-zero so a CI job stops. Check the output, fix the model or the test, deploy again. ```sh $ zsql deploy deploy tpcds (6240 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies PASSED store net paid by category FAILED catalog quantity by month --- generated SELECT DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", sum(T0."cs_quantity") AS "Catalog Quantity" FROM catalog_sales T0 JOIN date_dim T1 ON T0.cs_sold_date_sk = T1.d_date_sk GROUP BY DATE_TRUNC('month', T1."d_date")::DATE PASSED date range tests: 2 passed, 1 failed Error: tests failed $ echo $? 1 ``` ### The production branch Deploying while `main` (or whatever `production_branch` names) is checked out updates what production reads. The output says so: `as tpcds/main (the production branch)`. On the service a project owner can mark the production branch protected; then deploying to it needs owner access, and anyone else gets `Forbidden: deploying to the protected production branch 'main' of tpcds needs an owner`. ### --watch ```sh $ zsql deploy --watch deploy tpcds (6214 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies tests: 3 passed, 0 failed watching /home/you/tpcds for changes (ctrl-c to stop) ``` Deploys once, then polls the project directory every 500 ms for a changed `.yml` or `.yaml` file and redeploys 200 ms after a change. Errors are printed and watching continues. Pair it with the repl in a second terminal to edit a model and query it as you go. ## Branches The service branch is the checked-out git branch. There is no separate "promote" step: a branch is deployed by checking it out and running `zsql deploy`, and `main` is updated by merging and deploying `main`. - `zsql deploy` on `feature/returns` creates or updates `tpcds/feature/returns`. Nothing else changes. - Each deployed branch is a complete, isolated model. An application names the branch it wants in the URL: `/projects/tpcds/branches/feature/returns/sql`. - A query key granted a project without a branch reads `production_branch`. A key granted `*` reads every branch; a key granted one branch name reads that one. - `--branch NAME` overrides the git branch. CI checkouts are often detached, so pass it there (see [CI/CD](#cicd)). Deploying from a dirty working tree deploys the files on disk, committed or not. ## zsql status The deployed branch's summary, as the service returns it. ```sh $ zsql status { "project": "tpcds", "branch": "main", "deployed_at": "2026-10-01 15:04", "datasources": 1, "tables": 6, "fields": 42, "joins": 5, "paths": 9, "policies": 1, "tests": 3, "warnings": [] } ``` A branch that was never deployed answers `NotFound: no deployment for project tpcds branch feature/returns`. ## zsql list One line per deployment the key can see, across projects and branches. ```sh $ zsql list tpcds main 6 tables, 42 fields, 1 policies, 3 tests tpcds feature/returns 7 tables, 48 fields, 1 policies, 3 tests ``` `zsql list` works outside a project directory too; it only needs a key and a server. ## zsql remove Removes the deployed branch from the service. The confirmation is a flag. ```sh $ zsql remove Error: this removes tpcds/feature/returns from https://app.0sql.io; add --yes to confirm $ zsql remove --yes { "removed": "tpcds/feature/returns" } ``` Removing a branch removes its deployment, not the project or the git branch. ## zsql test Runs the tests the branch was deployed with, on the service, without redeploying. Useful after a key or policy change, and as a cheap health check of a branch. ```sh $ zsql test PASSED store net paid by category PASSED catalog quantity by month PASSED date range tests: 3 passed, 0 failed ``` Exits non-zero with `tests failed` when any fail. The format is on [Tests](#tests). ## Next steps - [Tests](#tests) - [CI/CD](#cicd) - [Querying from the CLI](#querying-from-the-cli) ### Modeling with zsql Model files are plain YAML you can write by hand. `zsql new` saves the typing: it writes a table or relation file with every key documented in comments, into the domain folder you name. `zsql check` then validates the whole project on the service and reports counts, warnings and test results without deploying anything. ## zsql new table ```sh zsql new table NAME --datasource DS [--physical-name P] [--domain core] ``` Writes `models//tbl..yml`, where the stem is the slug of `NAME` with `-` turned into `_`. `--physical-name` is the warehouse table; it defaults to the stem. An existing file is refused. ```sh $ zsql new table "Store Sales" --datasource tpcds --physical-name store_sales --domain store create models/store/tbl.store_sales.yml ``` The template, with the name, physical name and datasource filled in: ```yaml # A table and the dimensions and measures it holds. Joins between tables # live in rel.*.yml files. name: Store Sales physical_name: store_sales # defaults to the stem datasource: tpcds # The planner prefers the cheapest table that answers a question. Dimension # tables low, big facts high. cost: 10 # For a snapshot table (inventory, balances), the date dimension it snapshots on. # snapshot: Date # Data the table holds, so the planner prefers it for filters inside the # range and avoids it for filters outside. Predicates: between, # greater_than, greater_than_or_equal_to, less_than, less_than_or_equal_to # (number or date dimensions) and in_list (any dimension). # partitions: # - dimension: Date # predicate: between # filter_value: 2y # filter_value_end: 1d # - dimension: Region # predicate: in_list # filter_value: US, EU fields: # - type: dimension # dimension | measure # name: Category # description: Product category # data_type: string # string | integer | bigint | decimal | date | date_time | boolean # synonyms: [dept] # tags: [] # policy triggers, e.g. [pii] # expression: # sql: category # primary_key: false # # - type: measure # name: Revenue # data_type: decimal # format: currency:2 # expression: # sql: sum(revenue) ``` Fill it in: set `cost` (dimension tables low, big facts high), uncomment and edit the fields. A dimension's `expression.sql` is a column or a scalar expression; a measure's holds an aggregate. The file for the running example becomes: ```yaml name: Store Sales physical_name: store_sales datasource: tpcds cost: 100 fields: - type: dimension name: Store Ticket number description: Actual ticket number of an order data_type: integer expression: primary_key: true sql: ss_ticket_number - type: measure name: Store Net Paid description: Net amount paid for sales data_type: decimal expression: sql: sum(ss_net_paid) - type: measure name: Store Net Profit tags: [secure-cat] description: Net profit from sales data_type: decimal expression: sql: sum(ss_net_profit) ``` Every field key (`hidden`, `synonyms`, `format`, `grains`, `exclusions`, `inclusions`, `extended_blend_group`, `snapshot`, ...) is documented in [Semantic model](#semantic-model). ## zsql new relation ```sh zsql new relation --datasource DS [--domain core] ``` Writes `models//rel..yml`. One relation file per datasource per domain is the usual shape; it holds every join between that datasource's tables that the domain introduces. ```sh $ zsql new relation --datasource tpcds --domain store create models/store/rel.store.yml ``` The template, verbatim: ```yaml # Joins between this datasource's tables, by the names in their table files. # Cardinality is required; many_to_one is the common shape (fact to dimension). datasource: tpcds # orders_customer: # left: Orders # right: Customers # sql: left.customer_id = right.id # cardinality: many_to_one # many_to_one | one_to_many | one_to_one # join: inner # inner | left | right # # allow_measure_expansion: false ``` Each join is a key with `left` and `right` table names (the `name:` in their table files), a `sql` of the form `left.column = right.column`, and a `cardinality`. Filled in for the example: ```yaml datasource: tpcds store_sales_sold_date: left: Store Sales right: Date sql: left.ss_sold_date_sk = right.d_date_sk cardinality: many_to_one store_sales_item: left: Store Sales right: Item sql: left.ss_item_sk = right.i_item_sk cardinality: many_to_one ``` ## Naming - **Title Case names.** `Store Net Paid`, `Item Category`, `Month of Year`. Names are what callers write in specs and what appears as column aliases in the SQL. - **One name, one concept.** Table names are unique in the project; field names should be unique too, or callers must disambiguate with `@d` / `@m` or a uid. Prefix fact-specific fields with the fact (`Store Quantity`, `Catalog Quantity`) and leave conformed dimensions plain (`Date`, `Item Category`). - **Uids derive from names.** `Store Net Paid` becomes `store-net-paid`. Renaming a field changes its uid; callers that reference uids notice. - **Synonyms** are extra names callers can use: `synonyms: [dept]` on `Item Category` lets `dept` resolve, and fuzzy matching corrects close misspellings. ## Domain folders `models/` is read recursively; any depth works. The convention `zsql new` follows is one folder per business domain: `models/common/` for conformed dimensions (`Date`, `Item`, `Store`), `models/store/` for store sales and its joins, `models/catalog/`, `models/inventory/`. The folder name has no meaning to the loader; it keeps related tables and their relation file together. ``` models/ ├── common/ │ ├── tbl.date_dim.yml │ └── tbl.item.yml ├── store/ │ ├── tbl.store_sales.yml │ └── rel.store.yml └── inventory/ ├── tbl.inventory.yml └── rel.inventory.yml ``` ## zsql check ```sh zsql check ``` Packs the project exactly as `zsql deploy` would and sends it to the validate route. The service loads it, forms the join routes, runs the tests, and answers with the same counts a deploy returns, but keeps nothing. Use it after every edit. ```sh $ zsql check validate tpcds (6214 bytes) as tpcds/main (the production branch) to https://app.0sql.io validated tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies warning: models/store/tbl.store_returns.yml: table Store Returns has no cost; 0 assumed PASSED store net paid by category PASSED catalog quantity by month PASSED date range tests: 3 passed, 0 failed ``` The response holds `datasources`, `tables`, `fields`, `joins`, `paths` (join routes formed), `policies`, `tests`, a `warnings` list, and `test_results` with `passed`, `failed` and one entry per test. A model that does not load fails with one message naming the file: ```sh $ zsql check validate tpcds (6198 bytes) as tpcds/main (the production branch) to https://app.0sql.io DeployError: Error in models/store/tbl.store_sales.yml: Expression errors for Store Net Paid: Sql measure should have an aggregation function ``` Every error has the form `Error in : `. Warnings do not fail the check: a table without `cost`, an unknown key (`table Store Sales: unknown key 'owner' is ignored`), an ambiguous join route. Read them; they describe choices the planner made for you. The complete list is in [Troubleshooting](#troubleshooting). ## Agents If an LLM writes your model files, point it at [/docs/llms.txt](#docsllmstxt): the whole guide, including every YAML key, as one markdown file. Have it run `zsql check` after each edit and feed the errors back. ## Next steps - [Semantic model reference](#semantic-model) - [Tests](#tests) - [Deploying](#deploying) ### Projects A project is a directory of YAML: one `project.yml`, one `datasources.yml`, model files under `models/`, optional `security.yml` and `tests/`. It lives in git. `zsql init` writes the skeleton; the first `zsql deploy` creates the project on the service. ## zsql init ```sh zsql init [DIR] [--name NAME] [--uid UID] ``` `DIR` defaults to the current directory. `--name` defaults to the directory name; `--uid` defaults to a slug of the name. A directory that already holds a `project.yml` is refused; otherwise existing files are kept and missing ones created. ```sh $ zsql init tpcds create project.yml create datasources.yml create security.yml create models/.keep create tests/.keep update .gitignore git init (branch main) next: describe your warehouses in datasources.yml, `zsql new table ...`, `zsql auth --api-key ...`, `zsql deploy` ``` The result: ``` tpcds/ ├── .gitignore ├── project.yml ├── datasources.yml ├── security.yml ├── models/ └── tests/ ``` If the directory has no `.git`, `zsql init` runs `git init -q -b main`, so the first checked-out branch matches the default `production_branch`. The `.gitignore` gains: ``` # zsql: local secrets .zsql .strata ``` `.zsql` holds your API key. It does not belong in git. ## project.yml The template, verbatim: ```yaml # zsql project. The uid is what the server knows the project as; keep it stable. name: tpcds uid: tpcds description: # The branch that production reads from. Deploys go to the checked-out git # branch, so this is the branch to merge into when a change is ready. production_branch: main # A self-hosted zsqld, if any. The hosted service is used otherwise; the # api key lives in .zsql, never here. # server: http://127.0.0.1:3699 ``` | Key | Required | Meaning | |---|---|---| | `name` | yes | Display name. Missing it fails the deploy: `Cannot create project: name is required in project.yml`. | | `uid` | no | What the service knows the project as, and the `{uid}` in every API path (`/projects/{uid}/branches/{branch}/sql`). Lower-cased. Defaults to the slug of `name`. Keep it stable: changing it makes a new project. | | `description` | no | Written by the template, ignored by the loader. | | `production_branch` | no | The branch production reads from. Default `main`. Query keys granted without a branch read this one. | | `server` | no | An alternative server URL, read by the CLI only. Leave it commented to use the hosted service. | ### uid rules A uid is a slug: lower-case letters, digits, `-` and `_`. The default comes from `name` by lower-casing, keeping `[a-z0-9_-]`, turning every other run of characters into one `-`, and trimming trailing `-`. `Store Analytics (EU)` becomes `store-analytics-eu`. A `--uid` you pass to `zsql init` is lower-cased as given. The service refuses a deploy whose uid is not a slug. ## How the service learns a project Nothing registers a project ahead of time. The first `zsql deploy` of a uid creates it on the service with the `name` and `production_branch` from `project.yml`, and the user whose personal key made the deploy becomes its owner. Later deploys need write access on the branch; deploying to a protected production branch needs an owner. Access levels and query-key grants are managed in the console, see [Accounts](#accounts-keys-and-access). ## Branches A deployed branch is named after the checked-out git branch. `zsql deploy` on `feature/returns` deploys `tpcds/feature/returns`; merging to `main` and deploying there updates `tpcds/main`. Branches are isolated deployments: an application queries one by name, and nothing on `feature/returns` changes what `main` answers. `production_branch` is only a label the service uses for query keys granted without an explicit branch, and for the `(the production branch)` note `zsql deploy` prints. Branch semantics in full are on [Deploying](#deploying). ## Directory layout | Path | Holds | |---|---| | `project.yml` | `name`, `uid`, `production_branch` | | `datasources.yml` | One warehouse per key: `name`, `adapter`, `tier`, connection fields. See [Datasources](#datasources). | | `security.yml` | Row-level security policies (`policies: []` to start). See [Security](#row-level-security). | | `models/**/tbl.*.yml` | One table with its dimensions and measures. Any depth under `models/`; the folder is a domain by convention. | | `models/**/rel.*.yml` | Joins between one datasource's tables. | | `tests/*.yml` | Planner assertions, one per file, run on every deploy. See [Tests](#tests). | | `.zsql` | API key and server. Mode 0600, gitignored, never deployed. | Model files are found by filename prefix (`tbl.`, `rel.`) and extension (`.yml` or `.yaml`), recursively, in sorted order. Everything the loader reads is listed in [Semantic model](#semantic-model). ## Next steps - [Datasources](#datasources) - [Modeling with zsql](#modeling-with-zsql) - [Deploying](#deploying) ### Querying from the CLI Every query command sends a spec and an optional security context to the deployed branch and prints the service's answer. Write the query as a one-line shorthand (`--expr`) or as a JSON file (`--spec`). The CLI parses the shorthand locally into a spec, so what the service receives is always a spec. ## The shorthand in one minute A line of comma-separated items. Each is read from its shape: ``` item category, month(date), store net paid desc, item category in (Books, Music) yoy(store net paid), date between 2024-01-01 and 2024-12-31 Share = sum([Store Net Paid]@m) / sum([Catalog Net Paid]@m) ``` - A **filter** has a predicate: `=`, `!=`, `>`, `>=`, `<`, `<=`, `in (...)`, `not in (...)`, `between a and b`, `contains`, `starts with`, `ends with` (each with `not`), `is null`, `is not null`, `top N [by measure]`. - A **calculation** is a formula holding a bracket reference (`[Name]@m`, `[Name]@d`, `[Alias]`), named with `Alias = formula` or `formula as Alias`. - Anything else is a **projection**: a field name or uid, optionally wrapped in decorator functions (`month(...)`, `yoy(...)`, `pct(...)`, `running_sum(...)`, `moving_avg(..., 7)`), with an optional `as Alias` and `asc`/`desc`. - `[Name]` or `"Name"` is always a field, whatever the name says. Names match without regard to case; add `@d` or `@m` when a name is both a dimension and a measure. The full grammar is on [Shorthand](#shorthand-expressions); the spec it produces is on [Query spec](#the-query-spec). ## Security context A branch with policies in `security.yml` refuses to plan without a context (`ContextRequired: this branch has security policies; a context is required to plan`). Put the context in a file and pass `--context`: ```json {"email": "ana@example.com", "groups": [{"name": "Men", "tags": ["cat:Men"]}]} ``` The shape is on [Security context](#security-context). The transcripts below use the TPC-DS example, which has one policy, so they all pass `--context user.json`. ## zsql sql ```sh zsql sql (--expr LINE | --spec FILE) [--context FILE] [--explain] [--json] ``` ```sh $ zsql sql --expr "item category, store net paid" --context user.json projection item category projection store net paid -- datasource: TPC-DS (Postgres) SELECT T1."i_category" AS "Item Category", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_category" ``` The `projection` / `filter` / `calculation` lines say how each shorthand item was read; they go to stderr, so `zsql sql ... > query.sql` captures only the datasource comment and the SQL. With a spec file (a top-level `spec` key is unwrapped if present): ```json { "projections": [ {"field": "Item Category"}, {"field": "Date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Store Net Paid", "order_by": "desc"} ], "filters": [{"field": "Item Category", "predicate": "in_list", "value": "Books, Music"}] } ``` ```sh $ zsql sql --spec monthly.json --context user.json -- datasource: TPC-DS (Postgres) SELECT T2."i_category" AS "Item Category", DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN date_dim T1 ON T0.ss_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ss_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('books', 'music') GROUP BY T2."i_category", DATE_TRUNC('month', T1."d_date")::DATE ORDER BY sum(T0."ss_net_paid") desc ``` ### Fuzzy corrections A field name that is close but not exact is corrected, and the correction is printed before the SQL (to stderr): ```sh $ zsql sql --expr "departmnt, store net paid" --context user.json departmnt → Department (~0.8) -- datasource: TPC-DS (Postgres) SELECT ... ``` The format is `term → Field Name (~score)`. Exact and synonym matches print nothing. ### --json `--json` prints the service's response unchanged: `sql`, `datasource`, `datasource_uid`, `adapter`, `corrections` when any, and for shorthand queries the `spec` the CLI sent. ```sh $ zsql sql --expr "year, store quantity" --context user.json --json { "sql": "SELECT\n\tT1.\"d_year\" AS \"Year\",\n\tsum(T0.\"ss_quantity\") AS \"Store Quantity\"\nFROM\n\tstore_sales T0\n\tJOIN date_dim T1\n\t\tON T0.ss_sold_date_sk = T1.d_date_sk\nGROUP BY\n\tT1.\"d_year\"", "datasource": "TPC-DS (Postgres)", "datasource_uid": "tpcds", "adapter": "postgres" } ``` ## zsql explain ```sh zsql explain (--expr LINE | --spec FILE) [--context FILE] [--json] ``` The SQL, then the plan: the node tree the planner built and where the microseconds went. `zsql sql --explain` is the same. ```sh $ zsql explain --expr "item category, store net paid, catalog net paid" --context user.json projection item category projection store net paid projection catalog net paid -- datasource: TPC-DS (Postgres) WITH ag... AS (...), ag... AS (...) SELECT ... FULL OUTER JOIN ... -- 3 nodes merge final (root) join full projections: Item Category, Store Net Paid, Catalog Net Paid aggregation ag... on store_sales (tpcds) group by paths: store-sales -> item projections: Item Category, Store Net Paid security: Item Category in (men) aggregation ag... on catalog_sales (tpcds) group by paths: catalog-sales -> item projections: Item Category, Catalog Net Paid -- server 214 us: parser 11 us, plan 203 us -- phases (us): ... ``` Each node is `kind alias on table (datasource)`, with `group by`, the join type, `[segment]` and a transform or purpose when present, then indented lists of paths, projections, filters, security filters and masks. The SQL is abbreviated above; each `ag...` is one aggregation CTE. Two facts sharing `Item Category` become two `aggregation` nodes under one `merge`. `--json` returns the raw `nodes`, `timings` and `phases`. See [Explain](#explain) for the node fields. ## zsql explore ```sh zsql explore (--expr LINE | --spec FILE) [--context FILE] [-q TERM] [--json] ``` Given a query, what can it take on without breaking: the dimensions it can drill into and the measures it can add. `-q` narrows by name, uid or synonym, exact containing matches first, then close ones with a score. ```sh $ zsql explore --expr "item category, store net paid" --context user.json -q date date dim Date date year dim Year date month-of-year dim Month of Year date date-id dim Date ID date -- can add 4 dimensions, 0 measures · server 96 us ``` Without `-q`, every reachable dimension and measure is listed. Hidden fields are excluded. The last column names the tables a field lives on. ## zsql spec ```sh zsql spec LINE... ``` Local. Prints the spec JSON a shorthand line makes, with the item descriptions, and calls no server. Use it to learn the spec shape or to generate a file for `--spec`. ```sh $ zsql spec "item category, store net paid desc, item category in (Books, Music)" projection item category projection store net paid desc filter item category in_list Books, Music { "projections": [ {"field": "item category"}, {"field": "store net paid", "order_by": "desc"} ], "filters": [ {"field": "item category", "predicate": "in_list", "value": "Books, Music"} ] } ``` ## zsql fields and zsql tables ```sh zsql fields [Q] [--hidden] zsql tables ``` ```sh $ zsql fields net store-net-paid msr Store Net Paid store-net-profit msr Store Net Profit store-net-including-returns msr Store Net Including Returns catalog-net-paid msr Catalog Net Paid $ zsql fields catgory item-category dim Item Category ~0.86 $ zsql fields --hidden | grep hidden store-ticket-number dim Store Ticket number (hidden) ``` Each line is `uid dim|msr name`, with `(hidden)` and a `~score` for fuzzy matches. Without `Q` every field is listed; hidden ones only with `--hidden`. ```sh $ zsql tables date date_dim tpcds cost 10 item item tpcds cost 10 store-sales store_sales tpcds cost 100 catalog-sales catalog_sales tpcds cost 100 inventory inventory tpcds cost 100 ``` `uid physical_name datasource cost N`. These are the two discovery routes, `GET .../fields` and `GET .../tables`; see [Discovery](#discovery-explore-fields-and-tables). ## The repl ```sh zsql repl [--context FILE] ``` `zsql` with no subcommand inside a project also starts it. The prompt is `zsql> `. A plain line is shorthand; a line starting with `{` is a spec; both go to the sql route and print the SQL with the round trip time. The context given at start applies to every query. ``` $ zsql repl --context user.json tpcds/main on https://app.0sql.io: 6 tables, 42 fields, 5 joins, 1 policies type a line of projections and filters; .help for the forms, .quit to leave zsql> item category, store net paid, item category in (Books, Music) projection item category projection store net paid filter item category in_list Books, Music -- datasource: TPC-DS (Postgres) SELECT T1."i_category" AS "Item Category", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('books', 'music') GROUP BY T1."i_category" -- round trip 38.41 ms zsql> .explore item category, store net paid ? quantity store-quantity msr Store Quantity store-sales catalog-quantity msr Catalog Quantity catalog-sales inventory-quantity-on-hand msr Inventory Quantity On Hand inventory -- can add 0 dimensions, 3 measures · server 102 us zsql> .branch feature/returns tpcds/feature/returns on https://app.0sql.io: 7 tables, 48 fields, 6 joins, 1 policies zsql> .quit ``` Dot-commands, as in sqlite: | Command | Does | |---|---| | `.help`, `.?` | The shorthand forms and this list | | `.tables` | As `zsql tables` | | `.fields [q]` | As `zsql fields` | | `.branch [name]` | Print the branch, or switch to another deployed branch | | `.spec ` | The spec a line makes, without planning | | `.json {...}` | Plan a spec written inline | | `.explain ` | As `zsql explain` | | `.explore [? term]` | As `zsql explore`; the term after ` ? ` narrows | | `.quit`, `.exit`, `.q` | Leave | `help`, `quit` and the like also work without the dot unless a field has that name. Errors from the service print as `error: : ` and the session continues. History is kept in `~/.zsql/history` across sessions. ## Next steps - [Shorthand grammar](#shorthand-expressions) - [Query spec](#the-query-spec) - [Security context](#security-context) - [Query errors](#errors-and-corrections) ### Command reference `zsql [GLOBAL FLAGS] [ARGS]`. With no command inside a project directory, `zsql` starts the [repl](#repl); outside one it prints help. `zsql --version` prints the version; `zsql --help` prints that command's flags. ## Global flags | Flag | Default | Notes | |---|---|---| | `--project DIR` | the directory holding `project.yml`, searched from the current directory upwards | Error without one: ``no project.yml here or above; run `zsql init` to start a project`` | | `--branch NAME` | the checked-out git branch of the project; `main` outside a project | The service branch every command targets | | `--server URL` | see [server resolution](#environment-variables) | Trailing `/` trimmed | | `--uid UID` | `uid` from `project.yml` | Required outside a project directory | ## Project commands ### `zsql init [DIR] [--name NAME] [--uid UID]` | Argument | Default | Notes | |---|---|---| | `DIR` | `.` | Created if missing. Refused if it holds a `project.yml`. | | `--name` | the directory name | `name:` in `project.yml` | | `--uid` | slug of the name | Lower-cased | Writes `project.yml`, `datasources.yml`, `security.yml`, `models/.keep`, `tests/.keep`; appends to `.gitignore`; runs `git init -q -b main` if there is no `.git`. Existing files are kept. ### `zsql new table NAME --datasource DS [--physical-name P] [--domain D]` | Argument | Default | Notes | |---|---|---| | `NAME` | required | `name:` of the table, Title Case | | `--datasource` | required | A key in `datasources.yml` | | `--physical-name` | the file stem | Warehouse table name | | `--domain` | `core` | Folder under `models/` | Writes `models//tbl..yml`; the stem is the slug of `NAME` with `-` as `_`. Refused if the file exists. ### `zsql new relation --datasource DS [--domain D]` | Argument | Default | |---|---| | `--datasource` | required | | `--domain` | `core` | Writes `models//rel..yml`. ### `zsql new test NAME` Writes `tests/.yml`. ## Service commands ### `zsql auth --api-key KEY [--server URL]` Writes `api_key:` and optionally `server:` into `.zsql` in the project directory (mode 0600). Prints `saved to `. ### `zsql check` Same as `zsql deploy --dry-run`. ### `zsql deploy [--dry-run] [--watch]` | Flag | Does | |---|---| | `--dry-run` | POST to the validate route instead; the service keeps nothing | | `--watch` | Deploy, then redeploy whenever a `.yml` or `.yaml` under the project changes (poll 500 ms, debounce 200 ms) | Exits non-zero on a rejected archive or any failed test. ### `zsql test` Runs the deployed branch's tests on the service. Exits non-zero with `tests failed`. ### `zsql status` Prints the deployed branch's summary as JSON. ### `zsql list` One line per deployment: `project branch N tables, N fields, N policies, N tests`. Works outside a project. ### `zsql remove --yes` Removes the deployed branch. Without `--yes`: `this removes / from ; add --yes to confirm`. ### `zsql health` Prints ` ok`. No key needed. ## Query commands The three planning commands share the query arguments: | Flag | Notes | |---|---| | `--spec FILE` | A JSON file; a top-level `spec` key is unwrapped. Conflicts with `--expr`. | | `--expr LINE` | A shorthand line, parsed locally into a spec | | `--context FILE` | A JSON security context | One of `--spec` or `--expr` is required. ### `zsql sql (--spec FILE | --expr LINE) [--context FILE] [--explain] [--json]` | Flag | Does | |---|---| | `--explain` | Same output as `zsql explain` | | `--json` | Print the service's response as JSON | Prints corrections (stderr), `-- datasource: `, the SQL. ### `zsql explain (--spec FILE | --expr LINE) [--context FILE] [--json]` Corrections, `-- datasource:`, SQL, blank line, `-- N nodes`, the node tree, `-- server N us: parser N us, plan N us`, `-- phases (us): ...`. ### `zsql explore (--spec FILE | --expr LINE) [--context FILE] [-q TERM] [--json]` | Flag | Does | |---|---| | `-q TERM` | Keep fields whose name, uid or synonyms contain the term, or come close | Lines of `uid dim|msr name tables [~score]`, then `-- can add N dimensions, N measures · server N us`. ### `zsql spec LINE...` Prints the spec JSON a shorthand line makes. Local; no server, no key. ### `zsql repl [--context FILE]` | Flag | Does | |---|---| | `--context FILE` | Security context applied to every query in the session | Dot-commands: `.help` `.?` `.tables` `.fields [q]` `.branch [name]` `.spec ` `.json {...}` `.explain ` `.explore [? term]` `.quit` `.exit` `.q`. History in `~/.zsql/history`. ### `zsql fields [Q] [--hidden]` | Argument | Does | |---|---| | `Q` | Filter text; containing matches first, then fuzzy with a score | | `--hidden` | Include hidden fields | Lines of `uid dim|msr name [(hidden)] [~score]`. ### `zsql tables` Lines of `uid physical_name datasource cost N`. ## Environment variables | Variable | Used for | |---|---| | `ZSQL_API_KEY` | The API key. Wins over every file. | | `ZSQL_SERVER` | The server URL. Second after `--server`. | | `HOME` | Locates `~/.zsql/config` and `~/.zsql/history` | Server resolution: `--server`, `ZSQL_SERVER`, `server:` in `.zsql` (or `.strata`), `server:` in `~/.zsql/config`, `server:` in `project.yml`, then `http://127.0.0.1:3699`. Key resolution: `ZSQL_API_KEY`, `api_key:` in `.zsql` (or `.strata`), `api_key:` in `~/.zsql/config`. ## Config files | File | Holds | Notes | |---|---|---| | `project.yml` | `name`, `uid`, `production_branch`, optional `server` | Committed. Found by walking up from the current directory. | | `.zsql` | `api_key`, `server` | Project directory. Written by `zsql auth`, mode 0600, gitignored. YAML. | | `.strata` | Same keys | Read only when `.zsql` is absent. | | `~/.zsql/config` | `api_key`, `server` | Account-wide fallback. YAML. | | `~/.zsql/history` | repl history | | | `datasources.yml` | Warehouses | Shipped with secret keys removed | ## Exit codes and errors `zsql` exits `0` on success and non-zero on any error. Errors print as `Error: `. A service error prints its class and message, `: `: ``` Unauthorized: an API key is required: Authorization: Bearer Forbidden: query keys are read only; deploy with your personal key NotFound: no deployment for project tpcds branch main ContextRequired: this branch has security policies; a context is required to plan DeployError: Error in models/store/tbl.store_sales.yml: Expression errors for Store Net Paid: Sql measure should have an aggregation function ``` `zsql deploy` and `zsql test` exit non-zero after a successful HTTP call when any test failed (`tests failed`). The complete error class table is on [Query errors](#errors-and-corrections) and in the [API reference](#api-reference). ### Tests A test projects a list of fields and asserts on the SQL the planner generates for them. Tests live in `tests/*.yml`, one per file, and run on the service on every `zsql deploy`, `zsql check` and `zsql test`. They catch a changed join route, a lost aggregate, a renamed column, before an application does. ## Format ```yaml name: store net paid by category projections: - Item Category - Store Net Paid assert_regex: sum\(.*ss_net_paid.*\) ``` | Key | Required | Meaning | |---|---|---| | `name` | yes | Shown in the output. | | `projections` | yes | One or more field names (uids and synonyms work too). Projected plainly: no decorators, no filters, no security context. | | `assert_sql` | one of | The whole expected statement. Compared after lower-casing both sides and collapsing whitespace. | | `assert_regex` | one of | A regular expression the generated SQL must match. Case-insensitive; `.` matches newlines. | A test may carry both assertions. There is no `filters`, `context` or `spec` key: a test is projections only, planned with no security context. Policies therefore do not fire in tests; test the model, not the policy. Missing pieces fail the deploy: `Test must have a name`, `Test must have at least one field in 'projections'`, `Test must have at least one assertion (assert_sql or assert_regex)`. ## zsql new test ```sh $ zsql new test "store net paid by category" create tests/store_net_paid_by_category.yml ``` The template, verbatim: ```yaml # A planner test: project these fields and compare the SQL. Runs on every deploy. name: store net paid by category projections: - Category - Revenue # assert_sql: | # SELECT ... assert_regex: SELECT ``` Replace the projections with your fields and tighten the assertion. ## A worked example The model: `Store Sales` (fact, `cost: 100`) joins `Item` (`cost: 10`) through `store_sales_item` on `ss_item_sk = i_item_sk`. `Store Net Paid` is `sum(ss_net_paid)`, `Item Category` is `i_category`. `tests/store_net_paid_by_category.yml`: ```yaml name: store net paid by category projections: - Item Category - Store Net Paid assert_sql: | SELECT T1."i_category" AS "Item Category", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_category" ``` ```sh $ zsql deploy deploy tpcds (6214 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies PASSED store net paid by category tests: 1 passed, 0 failed ``` Now someone changes the join to `join: left`. The next deploy: ```sh $ zsql deploy deploy tpcds (6218 bytes) as tpcds/main (the production branch) to https://app.0sql.io deployed tpcds/main: 6 tables, 42 fields, 5 joins, 9 paths, 1 policies FAILED store net paid by category --- generated SELECT T1."i_category" AS "Item Category", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 LEFT JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_category" --- expected SELECT T1."i_category" AS "Item Category", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_category" tests: 0 passed, 1 failed Error: tests failed ``` The branch is deployed (the HTTP call succeeded), the CLI exits non-zero, and the diff is on screen. ## Output Each test prints as `PASSED`, `FAILED` or `ERROR` followed by its name. A failed `assert_sql` prints `--- generated` and `--- expected` blocks; a failed `assert_regex` prints the generated SQL (the response also carries the `pattern`). `ERROR` means the planner could not plan the projections at all, or the regex did not compile (`bad assert_regex: ...`); its message is printed under the name. The last line is `tests: N passed, N failed`. On the HTTP side the same data is `test_results` in the deploy response and the body of `POST /projects/{uid}/branches/{branch}/test`: ```json {"passed": 0, "failed": 1, "results": [ {"name": "store net paid by category", "status": "failed", "file": "store_net_paid_by_category.yml", "expected_sql": "SELECT T1...", "generated_sql": "SELECT T1..."} ]} ``` ## zsql test ```sh $ zsql test PASSED store net paid by category PASSED catalog quantity by month PASSED date range tests: 3 passed, 0 failed ``` Runs the tests stored with the deployed branch, on the service, without sending the project again. Exits non-zero with `tests failed` when any fail. ## Tips - **Prefer `assert_regex` for the thing you care about.** `sum\(.*ss_net_paid.*\)` survives alias changes, a reordered `GROUP BY` and a switch of adapter; a full `assert_sql` breaks on every one of them. Use `assert_sql` when the whole statement is the contract. - **Assert the join, not the whitespace.** `JOIN item` or `FROM\s+store_sales T0\s+JOIN item` pins the route; whitespace is already normalised for `assert_sql` and `.` matches newlines in `assert_regex`. - **One concern per file.** A test per fact-to-dimension route and one per compound measure keeps failures readable. - **Use the repl to write them.** `zsql sql --expr "item category, store net paid"` prints the SQL to paste into `assert_sql`. ## Next steps - [Deploying](#deploying) - [CI/CD](#cicd) - [Querying from the CLI](#querying-from-the-cli) --- ## Semantic Model How a 0sql project describes tables, fields, expressions and relationships, and what the planner does with them. ## Overview The semantic model is the core of 0sql. It describes warehouse tables as named dimensions and measures with joins between them, so that a query spec naming fields can be planned into one SQL statement for your warehouse. ## Critical Rules Before diving into components, understand these non-negotiable design principles: **tip: One Name, One Concept** **A field name identifies a single concept across your entire semantic layer.** There are no table-qualified names like `table.field_name`. Define the same name on more than one table and 0sql treats every definition as that one concept. "Total Revenue" on `store_sales` and "Total Revenue" on `catalog_sales` are both valid ways to answer a query for Total Revenue. At query time the planner picks the table whose reachable dimensions cover what the query asks for, and breaks ties by table `cost` (lower wins), so you can steer the choice. The rule therefore cuts the other way: **distinct concepts need distinct names.** If "Country" means the caller's country on one table and the ship-to country on another, name them "Caller Country" and "Ship Country", or 0sql will merge them. **Benefits:** - **Unambiguous references**: "Total Revenue" means exactly one thing, however many tables can produce it - **Compound measures**: `[Total Revenue] - [Total Cost]` works without table prefixes - **Conformed dimensions by naming**: a shared "Date" or "Customer" name is what lets tables blend - **Simple specs**: a request names fields, never tables, so callers and agents need no disambiguation **Considerations:** - Names are scoped per type: "Days To Recontact" can be both a dimension (the bucket) and a measure (the average), since a query asks for one or the other - Naming is where you apply domain knowledge: same thing, same name; different thing, different name ```yaml # In tbl.store_sales.yml - name: Total Revenue # one concept... # In tbl.catalog_sales.yml - name: Total Revenue # ...defined on a second table: either can serve it, # chosen by the dimensions requested, then by cost - name: Catalog Revenue # a different concept gets a different name ``` **danger: No Many-to-Many Relationships** **`many_to_many` cardinality is not supported.** This design constraint prevents double-counting bugs and ensures predictable aggregations. **Workaround:** Use junction tables with two separate relationships. ```yaml # ✗ NOT SUPPORTED users_roles: cardinality: many_to_many # ✓ USE JUNCTION TABLE users_user_roles: left: Users right: User Roles cardinality: one_to_many user_roles_roles: left: User Roles right: Roles cardinality: many_to_one ``` **Why:** This approach makes relationships explicit and prevents common aggregation errors where measures get counted multiple times. ## Key Components ### Tables Tables define semantic metadata for physical database tables: - **Dimensions**: Categorical fields for grouping - **Measures**: Aggregatable quantities - **Cost**: Query optimization hints - **Partitions**: Data availability constraints [Learn about tables →](#tables) ### Fields Fields are the building blocks of your semantic model: - **Dimensions**: Attributes for grouping and filtering - **Measures**: Quantities to aggregate - **Data Types**: string, integer, decimal, date, etc. - **Field metadata**: descriptions, synonyms and tags, returned by the fields endpoint; `format` and `display_type`, stored for your client [Learn about fields →](#dimensions) ### Expressions Expressions define how fields query the database: - **SQL**: Direct SQL expressions - **Lookups**: Dimension field indicators - **Arrays**: Multi-value fields - **Primary Keys**: Unique identifiers [Learn about expressions →](#sql-expressions) ### Relationships Relationships define how tables join: - **Cardinality**: one-to-one, one-to-many, many-to-one - **Join Types**: inner (default), left, right - **Join Conditions**: SQL expressions [Learn about relationships →](#cardinality) ## How It Works 1. **Model Definition**: you define tables, fields and relationships in YAML 2. **Deployment**: `zsql deploy` validates the project and ships it to 0sql 3. **Universe Formation**: the planner precomputes every join path from every table 4. **Query Planning**: your application POSTs a query spec and a security context; the planner picks tables and paths and writes the SQL 5. **SQL Returned**: the response is one SQL statement in the datasource's dialect, plus the datasource it targets 6. **Your application runs it**: 0sql never connects to the warehouse ## Advanced (in-depth and extra) The **Advanced** subsection below covers deeper and optional semantic-model features: exclusions, inclusions, semantic routing, and cost optimization. The core basics (tables, fields, expressions, relationships, imports) are above; Advanced is for more in-depth and extra scenarios. - **Imports**: Reuse field definitions across tables - **Partitions**: Define data availability constraints - **Snapshot Measures**: Point-in-time analysis - **Exclusions**: Filter-based measure exclusions - **Inclusions**: Multi-level calculations - **Decorators**: Temporal, window and contribution transforms applied in the query spec; see [Projections](#projections-and-decorators) [Advanced features →](#advanced-features) ## Next Steps - Read about [tables](#tables) - Explore [dimensions](#dimensions) and [measures](#measures) - Learn about [expressions](#sql-expressions) - Understand [relationships](#cardinality) - Go deeper in [Advanced](#advanced-features) Measures are the aggregated quantities of the model: what gets summed, counted, or averaged when a query groups by dimensions. ## Overview A measure is a field whose expression aggregates rows. The spec names dimensions, 0sql groups by them, and every measure on the query is computed over each group. Measures never appear in a `GROUP BY`; that is what separates them from [dimensions](#dimensions). Every measure resolves to a source table, that table's grain, and the dimensions reachable from it. That is what lets 0sql decide, per query, which dimensions a measure can be grouped by, which table should serve it, and how to blend it with measures from other facts without double counting. A measure carries this by construction; nothing about it is configured per combination. ## Measure Types Five kinds of measure are modeled in YAML. They share every property below and differ in how the expression is written and how the planner evaluates it. | Type | What it is | Declared by | Reference | |---|---|---|---| | **Standard** | One aggregate over the table's columns: `sum(amount)`, `count(distinct customer_id)` | `expression.sql` with an aggregate | this page | | **Compound** | A formula over other measures and dimensions by name: `[Total Revenue] - [Total Cost]`. Measures may come from different fact domains in the same datasource; components that can't reach a query dimension are auto-leveled | `expression.sql` with `[Name]` references | [Compound measures](#compound-measures) | | **Snapshot** | The value at the start or end of each period, for balances and inventory that must not be summed across time | `snapshot: beginning` or `ending`, on a table with `snapshot: ` | [Snapshot measures](#snapshot-measures) | | **Exclusion** (LOD) | Ignores chosen dimensions in its grouping or filtering, for percent of total, fixed denominators, hierarchy-wide totals | `exclusion_type` and `exclusions` | [Exclusions](#exclusions) | | **Inclusion** (LOD) | Aggregates at a finer level first, then re-aggregates at the query grain, for medians and averages of per-entity values | `inclusions` | [Inclusions](#inclusions) | Three more things can change what a measure returns, none of them modeled in YAML. They are request-time constructs written into the query spec and are covered under the Query API: - **[Calculations](#calculations)**: a formula the caller writes over the query's measures, such as a rate or a share. - **[Decorated projections](#projections-and-decorators)**: year-over-year, moving average, percent of total, and other transforms on a projected measure. - **[Segments](#segments)**: restrict a single measure to a cohort so it sits beside its unrestricted baseline. ## Definition ```yaml - type: measure name: Total Revenue description: Sum of all order amounts data_type: decimal format: currency:2 expression: sql: sum(amount) ``` ## Aggregation The aggregate in `expression.sql` is passed straight through into the generated SQL. 0sql does not restrict you to a fixed set of functions or rewrite them: whatever aggregate the warehouse supports, in the warehouse's own dialect, is a valid measure expression. `sum`, `avg`, `count`, `min`, and `max` work everywhere. Beyond those, use what the engine offers: `approx_count_distinct` or `hll` sketches on Snowflake, `percentile_cont` and `median`, `stddev` and `variance`, `listagg` or `string_agg`, `bool_or`, `any_value`, `uniqExact` and `quantile` on ClickHouse, and so on. ```yaml - type: measure name: Median Order Value data_type: decimal expression: sql: percentile_cont(0.5) within group (order by amount) - type: measure name: Unique Visitors (approx) data_type: integer expression: sql: approx_count_distinct(visitor_id) ``` Two consequences of pass-through: - **The expression is dialect-specific.** A measure is defined on one table in one datasource, so write it for that engine. When the same measure name is defined on a hot-tier table in a different engine, write that table's expression in its dialect; the shared name is what makes them one measure (see [One name, one concept](#semantic-model)). - **The planner still needs to recognize an aggregate.** 0sql knows each adapter's aggregate functions and uses that to tell a measure expression from a row-level one, which matters for grouping, for blending, and for [calculations](#calculations) written on top of measures. Standard SQL aggregates and each engine's own are recognized; a measure that references columns must contain a function call, or `zsql check` fails with `Sql measure should have an aggregation function`. Conditional aggregation is plain SQL: ```yaml - type: measure name: Completed Orders data_type: integer expression: sql: count(case when status = 'completed' then 1 end) ``` ## Properties | Property | Required | Description | |---|---|---| | `type` | yes | `measure` | | `name` | yes | The name specs and other fields reference. One name is one concept across the model; see [naming](#semantic-model) | | `data_type` | yes | The aggregate's result type, usually `decimal` or `integer`; a `min(order_date)` is a `date`. See [data types](#data-types) | | `expression.sql` | yes | The aggregate expression, in the datasource's dialect. Compound measures reference other fields as `[Name]` | | `description` | no | Returned by the fields endpoint, so callers and agents know what the measure means | | `format` | no | Presentation hint for your client, as a shortcut (`currency:2`, `percent:1`) or a mapping; nothing in SQL generation uses it. See [field format](#field-format) | | `hidden` | no | Hide from discovery (the fields endpoint omits it by default) while keeping it resolvable in a spec and available to compound measures and calculations | | `synonyms` | no | Other names people use for it ("gross revenue", "sales amount"); a spec reference resolves by name or synonym, and `fields?q=` matches them | | `snapshot` | no | `beginning` or `ending`; makes this a [snapshot measure](#snapshot-measures) | | `exclusion_type`, `exclusions` | no | Dimensions to leave out of grouping and/or filtering; see [exclusions](#exclusions) | | `inclusions` | no | Dimensions to aggregate by first; see [inclusions](#inclusions) | | `tags` | no | Free-form labels; [security policies](#row-level-security) trigger on them | ## Patterns **Ratio inside one table.** Both sides aggregate the same rows, so a single expression is enough: ```yaml - type: measure name: Profit Margin data_type: decimal format: percent:1 expression: sql: sum(profit) / nullif(sum(revenue), 0) ``` **Ratio across measures.** When the numerator and denominator are already measures, possibly on different tables, reference them by name so the planner blends them correctly: ```yaml - type: measure name: Revenue per Order data_type: decimal format: currency:2 expression: sql: "[Total Revenue] / nullif([Order Count], 0)" ``` **Share of total.** A measure that ignores the query's grouping for its denominator is an exclusion measure, not a ratio. See [exclusions](#exclusions). For an ad-hoc share on one request, the percent-of-total [decorator](#projections-and-decorators) needs no modeling. **Balances and inventory.** Summing a daily balance across a month is wrong; use a [snapshot measure](#snapshot-measures). ## Guidelines 1. **One aggregate per standard measure.** Put math between measures in a compound measure, where the planner can blend it, rather than in one expression that assumes one table. 2. **Guard denominators** with `nullif(..., 0)`. 3. **Match `data_type` to the result**, and let `format` handle display. 4. **Write descriptions and synonyms.** Synonyms resolve field references in specs and in `fields?q=` search; descriptions come back with the field so callers know what it means. 5. **Same concept, same name.** Define "Total Revenue" on every table that can serve it and let the planner choose; give a different concept a different name. 6. **Test with `zsql test`**: assert the SQL a projection produces, on every deploy. ## Next Steps - [Compound measures](#compound-measures) - [Snapshot measures](#snapshot-measures) - [Exclusions](#exclusions) and [Inclusions](#inclusions) - [Dimensions](#dimensions) - [Expressions](#sql-expressions) Fields are the building blocks of your semantic model: dimensions, measures, data types, and field format. ## Overview - **[Dimensions](#dimensions)**: Categorical attributes for grouping and filtering - **[Measures](#measures)**: Aggregatable quantities (sum, count, avg, etc.) - **[Data Types](#data-types)**: `string`, `integer`, `decimal`, `date`, `date_time`, `boolean`, and more - **[Field format](#field-format)**: currency, percent and date formats, stored as metadata for your client ## Quick Reference | Field type | Use in GROUP BY | Aggregation | Example | |------------|-----------------|-------------|---------| | Dimension | Yes | No | Customer Name, Region, Order Date | | Measure | No | Yes (required) | Total Revenue, Order Count | ## Next Steps - [Dimensions](#dimensions) - [Measures](#measures) - [Data types](#data-types) - [Expressions](#sql-expressions): How fields query the database Dimension fields represent categorical attributes used for grouping and filtering. ## Overview Dimensions are the categorical fields in your semantic model. They represent attributes of your data that can be used for grouping, filtering, and analysis. ## Characteristics - **Used in GROUP BY** - Dimensions group query results - **Can be filtered** - Specs filter by dimension values - **Lower cardinality** - Typically fewer unique values than measures - **Examples**: Customer Name, Product Category, Order Date, Region ## Basic Definition ```yaml - type: dimension name: Customer Name description: Name of the customer data_type: string expression: sql: customer_name ``` ## Dimension Properties ### name The name a spec uses to reference the dimension (case-insensitive), and the name your client sees. ```yaml name: Customer Name ``` ### description Optional description explaining what the dimension represents. ```yaml description: Full name of the customer ``` ### data_type The data type of the dimension. See [data types](#data-types). ```yaml data_type: string ``` For **date and date_time** dimensions, including grains and filtering patterns, see [Date and DateTime dimensions](#date-and-datetime-dimensions). ### expression How to query this dimension from the database. ```yaml expression: sql: customer_name ``` [Learn about expressions →](#sql-expressions) ## Optional Properties ### hidden Hide the dimension from discovery (the fields endpoint omits it by default, `explore` never suggests it) while keeping it resolvable in a spec and usable in complex dimensions and calculations. ```yaml hidden: true ``` ### display_type Tell your client what kind of value this is. Stored as metadata; nothing in SQL generation uses it. ```yaml display_type: url # default, html, url, email, phone_number, image ``` ### format Presentation hint for your client (shortcut string or mapping). Validated and stored at deploy; nothing in SQL generation uses it. ```yaml format: currency:2 # or hash: { type: currency, precision: 2 } ``` [Learn about field format →](#field-format) ### synonyms Alternative names for this dimension. A field reference in a spec resolves by uid, name or synonym, and `GET .../fields?q=` matches on them, so callers using different terminology land on the same field. ```yaml synonyms: - client name - buyer name - account name ``` ### extended_blend_group For dimensions that should blend across tables in queries, set the same `extended_blend_group` on each. See [Extended blending groups](#extended-blending-groups). ## Primary Keys Mark dimensions as primary keys for query optimization: ```yaml - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id ``` **Benefits:** - Helps planner optimize joins - Ensures uniqueness - Improves query performance ## Complex Dimensions Complex dimensions are dimensions whose expressions are defined in terms of **other dimensions** instead of directly referencing a physical column. Rules: - Complex dimensions may reference **only dimensions**, not measures. - All referenced dimensions must be reachable in the same **universe** as the complex dimension (so they can be grouped together). - Evaluation happens at query time, similar to [compound measures](#compound-measures). ### Example: Derived status label ```yaml - type: dimension name: Customer Status Code data_type: string expression: sql: status_code - type: dimension name: Customer Status data_type: string expression: sql: | CASE WHEN [Customer Status Code] = 'A' THEN 'Active' WHEN [Customer Status Code] = 'I' THEN 'Inactive' ELSE 'Unknown' END ``` ### Example: Custom date bucket ```yaml - type: dimension name: Order Date data_type: date expression: sql: order_date - type: dimension name: Order Month Label data_type: string expression: sql: to_char([Order Date], 'YYYY-MM') ``` Use complex dimensions when you want to: - Reuse existing dimensions in multiple derived attributes, - Keep business-friendly labels and buckets out of raw SQL in every query, and - Ensure that all logic for a concept (like “status” or “month”) lives in one place in the semantic model. ## Common Patterns ### Simple Dimension ```yaml - type: dimension name: Product Category data_type: string expression: sql: category_name ``` ### Formatted Dimension ```yaml - type: dimension name: Product Image data_type: string display_type: image expression: sql: image_url ``` ### Primary Key Dimension ```yaml - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id ``` ## Best Practices 1. **Use descriptive names** - Business-friendly dimension names 2. **Add descriptions** - They come back from the fields endpoint, so callers know what a dimension means 3. **Set grains for date/datetime** - See [Date and DateTime dimensions](#date-and-datetime-dimensions) 4. **Mark primary keys** - Improve query optimization 5. **Use appropriate data types** - Match database column types ## Next Steps - [Date and DateTime dimensions](#date-and-datetime-dimensions) - Grains, formatting, and date-specific patterns - [Measures](#measures) - [Data types](#data-types) - [Field format](#field-format) - [Expressions](#sql-expressions) 0sql emits SQL for 13 warehouse dialects. The `adapter` on a datasource picks which one. ## Supported Adapters | Identifier | Engine | Identifier quoting | |---|---|---| | [`athena`](#athena-adapter) | AWS Athena | `"double quotes"` | | [`bigquery`](#bigquery-adapter) | Google BigQuery | `` `backticks` `` | | [`clickhouse`](#clickhouse-adapter) | ClickHouse | `` `backticks` `` | | [`databricks`](#databricks-adapter) | Databricks SQL | `` `backticks` `` | | [`druid`](#druid-adapter) | Apache Druid | `"double quotes"` | | [`duckdb`](#duckdb-adapter) | DuckDB | `"double quotes"` | | [`mysql`](#mysql-adapter) | MySQL | `` `backticks` `` | | [`postgres`](#postgresql-adapter) | PostgreSQL | `"double quotes"` | | [`redshift`](#redshift-adapter) | Amazon Redshift | `"double quotes"` | | [`snowflake`](#snowflake-adapter) | Snowflake | `"double quotes"` | | [`sqlite`](#sqlite-adapter) | SQLite | `"double quotes"` | | [`sqlserver`](#sql-server-adapter) | Microsoft SQL Server, Azure SQL | `[brackets]` | | [`trino`](#trino-adapter) | Trino, Presto | `"double quotes"` | Any other value fails validation with `adapter '…' is not supported`. ## What an adapter controls An adapter is a dialect. It decides the shape of the SQL you get back from `/sql`: - **Identifier quoting**, as in the table above. Quoted identifiers in an `expression.sql` keep their case; bare words are lowercased. - **Date functions**: how a `truncate` decorator becomes `DATE_TRUNC('month', x)`, `DATE_TRUNC(x, MONTH)` or `DATETRUNC(month, x)`, how year, month and day-of-week are extracted, and how a date literal is written. - **Join support for blends**: when a query mixes measures from two facts, the per-fact subqueries are combined with a `FULL OUTER JOIN` on their shared dimensions. MySQL has no `FULL JOIN`, so there the blend is a `LEFT JOIN`. - **Subquery structure**: CTEs everywhere except Druid, which gets nested subqueries. - **Which bare words are columns**: each dialect's reserved keywords and aggregate functions decide what in an expression is a column reference and what counts as an aggregate. There are no adapter-specific options to set in the project. Per-query overrides exist through the spec's `db_settings`, not through YAML. ## Common Configuration Every datasource takes the same keys: - **adapter** (required), one of the identifiers above - **name**, display name (defaults to the key) - **tier**, `hot`, `warm` or `cold`; see [Datasources](#datasources) - **description** - **query_timeout** and **extra_query_params**, carried as metadata for the caller - connection fields, carried as metadata ## Connection fields and secrets 0sql never connects to your warehouse, so the connection fields on a datasource are carried as metadata for your own application. They are not validated, and 0sql never uses them. Keep secrets out of the project anyway. If `password`, `private_key`, `personal_access_token`, `secret_access_key`, `access_key_id`, `oauth_client_id`, `oauth_client_secret` or `api_key` do land in `datasources.yml`, `zsql deploy` strips them before building the archive. ## Next Steps - Read the adapter page for your warehouse - [Datasources](#datasources), tiers and full configuration - [Managing Datasources](#datasources) ### Athena Adapter Generate Athena SQL from your model. Athena is the usual `cold` tier: history in S3 behind a hotter warehouse. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date literals as `DATE '2026-01-01'`; grain truncation as `DATE_TRUNC('month', x)`; parts with `YEAR(x)`, `DAY_OF_WEEK(x)`, `DAY_OF_YEAR(x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml athena_data: adapter: athena name: AWS Athena tier: cold region: us-east-1 database: analytics catalog: awsdatacatalog workgroup: primary # access_key_id / secret_access_key: never here; your application holds them ``` ## Connection fields 0sql never connects to Athena. Region, database, catalog, workgroup and any S3 output location are carried as metadata for your own application, which runs the SQL it gets back with its own AWS credentials (an IAM role on your compute, or keys). 0sql never uses them. If keys do land in `datasources.yml`, `zsql deploy` strips `access_key_id` and `secret_access_key` before building the archive. ## Notes - **Expressions are Athena SQL.** `approx_distinct`, `arbitrary`, `array_agg` and `date_diff` are fine in `expression.sql`. - **Pair with a hot tier.** Give the Athena tables a date [partition](#partitions) covering the history they hold and a higher `cost`, so the planner routes recent-data queries to the hot tier and only history to Athena. See [Semantic Routing](#semantic-routing). ## Next Steps - [Datasources](#datasources) - [Semantic Routing](#semantic-routing) ### BigQuery Adapter Generate BigQuery Standard SQL from your model. ## What the adapter emits - Identifiers in `` `backticks` ``; bare column names in expressions are lowercased. - Date literals as `DATE '2026-01-01'`; grain truncation as `DATE_TRUNC(x, MONTH)` (BigQuery's argument order); parts with `EXTRACT(YEAR FROM x)`, `EXTRACT(DAYOFWEEK FROM x)`. - CTEs for the per-fact subqueries (no temp tables) and a `FULL OUTER JOIN` when a query blends measures from two facts. - BigQuery's own aggregates are recognized as aggregates in measure expressions: `any_value`, `countif`, `logical_and`, `logical_or`, `array_concat_agg`, `approx_quantiles`, `approx_top_count`, `approx_top_sum`, on top of the standard list. The same monthly revenue projection as on the [Postgres page](#postgresql-adapter), in BigQuery's dialect: ```sql SELECT DATE_TRUNC(T1.`d_date`, MONTH) AS `Month(Date)`, sum(T0.`ss_net_paid`) AS `Store Net Paid` FROM `analytics.tpcds.store_sales` T0 JOIN `analytics.tpcds.date_dim` T1 ON T0.ss_sold_date_sk = T1.d_date_sk GROUP BY DATE_TRUNC(T1.`d_date`, MONTH) ``` ## Configuration ```yaml bq: adapter: bigquery name: BigQuery tier: warm project: my-gcp-project dataset: analytics location: US ``` ## Connection fields 0sql never connects to BigQuery. `project`, `dataset` and `location` are carried as metadata for your own application, which runs the SQL it gets back with its own client and credentials. 0sql never uses them. Keep service-account keys out of the project entirely; if a key does land in `datasources.yml`, `zsql deploy` strips `private_key` and `api_key` before building the archive. ## Notes - **Qualify tables in `physical_name`.** Write `project.dataset.table` or `dataset.table`; the value is emitted verbatim in `FROM`. - **Expressions are BigQuery SQL.** `countif(status = 'done')`, `approx_count_distinct(user_id)` and `SAFE_DIVIDE` are fine in `expression.sql`. ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### ClickHouse Adapter Generate ClickHouse SQL from your model, typically for a hot tier in front of the warehouse. ## What the adapter emits - Identifiers in `` `backticks` ``; bare column names in expressions are lowercased. - Date literals as `toDate('2026-01-01')`; grain truncation as `date_trunc('month', x)`; parts with `toYear(x)`, `toMonth(x)`, `toDayOfWeek(x)`. - CTEs for the per-fact subqueries (no temp tables), window functions where a decorator needs them, and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml clickhouse_hot: adapter: clickhouse name: ClickHouse Hot Tier tier: hot protocol: https host: abc123.us-east-1.aws.clickhouse.cloud port: 8443 database: analytics username: analyst # password: never here; zsql deploy strips it if present ``` ## Connection fields 0sql never connects to ClickHouse. Protocol, host, port and database are carried as metadata for your own application, which runs the SQL it gets back over whichever interface it prefers. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Expressions are ClickHouse SQL.** `uniqExact`, `quantile(0.5)(x)`, `argMax` and `countIf` are fine in `expression.sql`. - **Built for the hot tier.** Give the ClickHouse tables a low `cost` and a date [partition](#partitions) covering what they hold, so the planner routes recent-data queries here and falls back to the warehouse otherwise. See [Semantic Routing](#semantic-routing). - **Views work.** A view or materialized view is referenced by `physical_name` like any table. ## Next Steps - [Datasources](#datasources) - [Semantic Routing](#semantic-routing), one model across a warehouse and a hot tier ### Databricks Adapter Generate Databricks SQL from your model. ## What the adapter emits - Identifiers in `` `backticks` ``; bare column names in expressions are lowercased. - Grain truncation as `DATE_TRUNC('month', x)`; parts with `YEAR(x)`, `DAYOFYEAR(x)`, `DAYOFWEEK(x)`, `WEEKOFYEAR(x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml databricks_lakehouse: adapter: databricks name: Databricks Lakehouse tier: warm host: workspace-id.cloud.databricks.com warehouse: 1234abcd5678ef90 catalog: main schema: default # oauth_client_id / oauth_client_secret: never here; your application holds them ``` ## Connection fields 0sql never connects to Databricks. Host, SQL warehouse id, catalog and schema are carried as metadata for your own application, which runs the SQL it gets back with its own service principal. 0sql never uses them. If secrets do land in `datasources.yml`, `zsql deploy` strips `oauth_client_id`, `oauth_client_secret`, `personal_access_token` and `password` before building the archive. ## Notes - **Expressions are Databricks SQL.** `approx_count_distinct`, `percentile_approx`, `collect_list` and `try_divide` are fine in `expression.sql`. - **Unity Catalog names** go in `physical_name` as `catalog.schema.table`, or `schema.table` when your application's session sets a default catalog. ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### Druid Adapter Generate Druid SQL from your model, with subqueries in place of CTEs. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - **No CTEs or temp tables.** The per-fact subqueries that other dialects put in a `WITH` clause are nested inline as subqueries; blends still use a `FULL OUTER JOIN`. - Date parts with `TIME_EXTRACT(x, 'YEAR')`, `TIME_EXTRACT(x, 'DOW')` and so on. - The end of a `between` date filter is extended to the last hour of the day, so a filter through `2026-01-31` includes that day's rows, and date filters are not pushed down greedily into every subquery. ## Configuration ```yaml druid_cluster: adapter: druid name: Druid Cluster tier: hot protocol: http host: druid.example.com port: 8082 ``` ## Connection fields 0sql never connects to Druid. Protocol, host and port are carried as metadata for your own application, which runs the SQL it gets back against the broker. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Expressions are Druid SQL.** `APPROX_COUNT_DISTINCT_DS_HLL`, `TIME_FLOOR` and `LATEST` are fine in `expression.sql`. - **A natural hot tier.** Druid usually holds recent, pre-aggregated data in front of a warehouse of record. Give its tables a low `cost` and a date [partition](#partitions) so the planner routes recent-data queries there. See [Semantic Routing](#semantic-routing). ## Next Steps - [Datasources](#datasources) - [Semantic Routing](#semantic-routing) ### DuckDB Adapter Generate DuckDB SQL from your model. DuckDB is the usual choice for local development and tests: one file, no server, and the same model deploys unchanged against a warehouse adapter later. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date literals as `DATE '2026-01-01'`; grain truncation as `DATE_TRUNC('month', x)`; parts with `EXTRACT('year' FROM x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml local: name: Local DuckDB adapter: duckdb tier: hot file: ./data/analytics.duckdb ``` ## Connection fields 0sql never opens the file. `file` is carried as metadata for your own application, which runs the SQL it gets back; 0sql never uses it. A relative path resolves against the project directory. There is nothing secret to keep out of git. ## Notes - **Expressions are DuckDB SQL.** `quantile_cont`, `list_aggregate`, `strftime` and `::` casts all work in `expression.sql`. - **Same model, two adapters.** A common layout is a `local` DuckDB datasource for development and a warehouse datasource for production, with the same tables defined on both. The planner picks by tier and cost; see [Semantic Routing](#semantic-routing). ## Next Steps - [Datasources](#datasources) - [Managing Datasources](#datasources) ### MySQL Adapter Generate MySQL SQL from your model. ## What the adapter emits - Identifiers in `` `backticks` ``; bare column names in expressions are lowercased. - **No `FULL JOIN`.** When a query blends measures from two facts, the per-fact subqueries are combined with a `LEFT JOIN` instead of the `FULL OUTER JOIN` other dialects get. Groups that exist only on the right-hand fact are not returned. - MySQL has no `DATE_TRUNC`, so grain truncation is built from MySQL's own date functions; parts with `YEAR(x)`, `MONTH(x)`, `DAYOFWEEK(x)`. - CTEs for the per-fact subqueries and window functions where a decorator needs them; both require MySQL 8.0 or later. ## Configuration ```yaml mysql_db: adapter: mysql name: MySQL Database tier: warm host: db.example.com port: 3306 database: analytics username: analyst # password: never here; zsql deploy strips it if present ``` ## Connection fields 0sql never connects to MySQL. The fields above are carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Expressions are MySQL SQL.** `group_concat`, `if()` and `date_format` are fine in `expression.sql`. - **Blends are left joins.** Put the fact whose groups must all appear first in the spec's projections, or model a conformed dimension table so each fact's measure can be served from a path that reaches it. ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### PostgreSQL Adapter Generate PostgreSQL SQL from your model. `adapter: postgres` is also the dialect the running TPC-DS example uses. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date literals as `'2026-01-01'::DATE`; grain truncation as `DATE_TRUNC('month', x)`; parts with `EXTRACT(year FROM x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. A monthly revenue projection comes back in this shape: ```sql SELECT DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", sum(T0."ss_net_paid") AS "Store Net Paid" FROM store_sales T0 JOIN date_dim T1 ON T0.ss_sold_date_sk = T1.d_date_sk GROUP BY DATE_TRUNC('month', T1."d_date")::DATE ``` ## Configuration ```yaml warehouse: adapter: postgres name: Production Warehouse tier: hot host: db.example.com port: 5432 # default 5432 database: analytics schema: public username: analyst # password: never here; zsql deploy strips it if present ``` ## Connection fields 0sql never connects to Postgres. The fields above are carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips secret keys from the file, so nothing sensitive reaches the service. ## Notes - **Expressions are Postgres SQL.** `percentile_cont(0.5) within group (order by x)`, `string_agg`, `::` casts and `date_part` are all fine in `expression.sql`. - **Schema-qualified tables** go in `physical_name` (`sales.orders`); it is emitted verbatim in `FROM`. ## Next Steps - [Datasources](#datasources), tiers and secrets - [Managing Datasources](#datasources) ### Redshift Adapter Generate Amazon Redshift SQL from your model. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date literals as `'2026-01-01'::DATE`; grain truncation as `DATE_TRUNC('month', x)`; parts with `EXTRACT(YEAR FROM x)`, `EXTRACT(DOW FROM x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. - No array functions: `expression.array: true` is not supported on this adapter. ## Configuration ```yaml redshift_dw: adapter: redshift name: Redshift Warehouse tier: warm host: default-workgroup.123456789012.us-east-1.redshift-serverless.amazonaws.com port: 5439 database: analytics schema: analytics username: reader # password: never here; zsql deploy strips it if present ``` ## Connection fields 0sql never connects to Redshift. The fields above are carried as metadata for your own application, which runs the SQL it gets back; provisioned clusters and serverless workgroups differ only in the endpoint your application dials. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Postgres-like, not Postgres.** Redshift speaks the Postgres wire protocol but its SQL differs; the adapter emits Redshift forms where they diverge. Write `expression.sql` for Redshift: `approximate count(distinct x)`, `listagg`, `median`. - **Pair with a hot tier.** Redshift is a good system of record; route recent-data queries to ClickHouse or DuckDB in front of it with a [partition](#partitions). See [Semantic Routing](#semantic-routing). ## Next Steps - [Datasources](#datasources) - [PostgreSQL adapter](#postgresql-adapter) ### Snowflake Adapter Generate Snowflake SQL from your model. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased, so quote a column that was created with a mixed-case name. - Date literals as `'2026-01-01'::DATE`; grain truncation as `DATE_TRUNC('month', x)`; parts with `YEAR(x)`, `MONTH(x)`, `DAYOFWEEK(x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml snowflake_dw: adapter: snowflake name: Snowflake Data Warehouse tier: hot account_identifier: myorg-myaccount database: ANALYTICS_DB warehouse: COMPUTE_WH schema: PUBLIC role: ANALYST username: analyst # personal_access_token or private_key: never here; your application holds them ``` ## Connection fields 0sql never connects to Snowflake. `account_identifier`, `database`, `warehouse`, `schema` and `role` are carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses them. If secrets do land in `datasources.yml`, `zsql deploy` strips `password`, `private_key`, `personal_access_token`, `oauth_client_id` and `oauth_client_secret` before building the archive. ## Notes - **Expressions are Snowflake SQL.** `approx_count_distinct`, `hll`, `median`, `listagg` and `iff` are fine in `expression.sql`. - **Qualify tables** in `physical_name` (`ANALYTICS_DB.PUBLIC.STORE_SALES`) if your application's session does not set a default database and schema. ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### SQLite Adapter Generate SQLite SQL from your model, for demos and small embedded datasets. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date parts with `strftime`: `CAST(strftime('%Y', x) AS INTEGER)`, `CAST(strftime('%w', x) AS INTEGER)` and so on. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts (SQLite 3.39 or later). ## Configuration ```yaml sqlite_local: adapter: sqlite name: Local SQLite tier: hot file: ./data/warehouse.db ``` ## Connection fields 0sql never opens the file. `file` is carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses it. There is nothing secret to keep out of git. ## Notes - **Expressions are SQLite SQL.** `total()`, `group_concat` and `strftime` are fine in `expression.sql`; there is no `percentile_cont`. - **Prefer DuckDB for analytics.** For anything beyond a demo, [DuckDB](#duckdb-adapter) has the analytical SQL (date truncation, window functions, quantiles) that the generated statements lean on. ## Next Steps - [Datasources](#datasources) - [DuckDB adapter](#duckdb-adapter) ### SQL Server Adapter Generate T-SQL for Microsoft SQL Server and Azure SQL from your model. ## What the adapter emits - Identifiers in `[brackets]`; bare column names in expressions are lowercased. - Grain truncation as `DATETRUNC(month, x)` (SQL Server 2022 and Azure SQL); parts with `YEAR(x)`, `DATEPART(QUARTER, x)`, `DATEPART(WEEKDAY, x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. - No array functions: `expression.array: true` is not supported on this adapter. ## Configuration ```yaml sqlserver_db: adapter: sqlserver name: SQL Server Database tier: warm host: server.database.windows.net port: 1433 database: analytics username: analyst # password: never here; zsql deploy strips it if present ``` Azure SQL uses the same adapter; the difference is in how your application connects, not in the SQL. ## Connection fields 0sql never connects to SQL Server. The fields above are carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Expressions are T-SQL.** `count_big`, `iif()`, `string_agg` and `try_cast` are fine in `expression.sql`. - **Schema-qualified tables** go in `physical_name` (`dbo.store_sales`). ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### Trino Adapter Generate Trino and Presto SQL from your model. ## What the adapter emits - Identifiers in `"double quotes"`; bare column names in expressions are lowercased. - Date literals as `DATE '2026-01-01'`; grain truncation as `DATE_TRUNC('month', x)`; parts with `YEAR(x)`, `DAY_OF_WEEK(x)`, `WEEK(x)`. - CTEs for the per-fact subqueries and a `FULL OUTER JOIN` when a query blends measures from two facts. ## Configuration ```yaml trino_cluster: adapter: trino name: Trino Cluster tier: warm host: trino.example.com port: 8080 catalog: hive schema: default username: analyst # password: never here; zsql deploy strips it if present ``` ## Connection fields 0sql never connects to Trino. Host, catalog and schema are carried as metadata for your own application, which runs the SQL it gets back. 0sql never uses them. If a password does land in `datasources.yml`, `zsql deploy` strips it before building the archive. ## Notes - **Expressions are Trino SQL.** `approx_distinct`, `approx_percentile`, `array_agg` and `date_diff` are fine in `expression.sql`. - **Qualify tables** in `physical_name` (`hive.default.store_sales`) when the application's session does not set a default catalog and schema. ## Next Steps - [Datasources](#datasources) - [Adapters](#adapters) ### Datasources Declare the warehouses your tables live in, pick a SQL dialect per datasource, and understand how tiers influence routing. ## What are Datasources? `datasources.yml` names every warehouse the project models. Each entry tells 0sql: - Which **adapter** to emit SQL for (postgres, snowflake, bigquery, ...). The adapter is the SQL dialect of the statement you get back. - The **tier** (`hot`, `warm`, `cold`), which influences which datasource the planner routes a query to. - Connection details (host, port, database, schema), carried as metadata for your application. 0sql never connects to your warehouse. It plans against the model and returns one SQL statement; your application runs it. See [Managing Datasources](#datasources). ## File format `datasources.yml` is a mapping of **datasource key** to settings. The key, lowercased, is the datasource uid that table files refer to in their `datasource:` line. This is the template `zsql init` writes: ```yaml # The warehouses this project's tables live in. One entry per datasource, # keyed by a short id that table files refer to. Connection details that # are secrets (password, tokens, keys) are stripped before deploy, and # nothing here is ever used by 0sql. # # warehouse: # name: Warehouse # adapter: postgres # duckdb | postgres | snowflake | bigquery | databricks | trino | mysql | sqlserver | ... # tier: hot # hot | warm | cold; the planner prefers hotter tiers # host: localhost # port: 5432 # database: analytics # username: analyst # schema: public # # local: # name: Local DuckDB # adapter: duckdb # tier: hot # file: ./data/analytics.duckdb ``` ### Keys | Key | Required | Meaning | |---|---|---| | `adapter` | yes | One of the 13 [adapter identifiers](#adapters). Anything else fails with `adapter '…' is not supported` | | `name` | no | Display name; defaults to the key | | `tier` | no | `hot`, `warm` or `cold` | | `description` | no | Free text | | `query_timeout`, `extra_query_params` | no | Accepted and carried as metadata for the caller; 0sql does not run queries, so it never applies them | | connection fields | no | `host`, `port`, `database`, `username`, `schema`, `file`, `account_identifier`, `warehouse`, `role`, ... Not validated; carried as metadata for your own application and never used by 0sql | The file is required, needs at least one entry, and refuses a repeated key. ## Tier Configuration 0sql has three tiers. They matter when the same table or measure is modeled in more than one datasource: the planner prefers hotter tiers when choosing which datasource's SQL to generate. See [Semantic Routing](#semantic-routing). ### hot The datasource you want queries to land on first: the main warehouse, or a fast OLAP tier in front of it. ```yaml warehouse: adapter: postgres tier: hot # ... connection details ``` ### warm A secondary datasource the planner falls back to when the hot tier cannot answer the whole query. ```yaml archive: adapter: postgres tier: warm # ... connection details ``` ### cold Archive or slow-access data, used last. Typical for object-store engines holding history. ```yaml historical: adapter: athena tier: cold # ... connection details ``` Tier is one of three routing inputs. Table `cost` and [partitions](#partitions) declared on tables are the other two: a partition says which slice of data a table holds, so the planner can route a filter on last month to the hot tier and a filter on 2019 to the cold one. ## Secrets Connection fields that are secrets never leave your machine: - 0sql never connects to your warehouse, so no credential belongs in the project at all. - `zsql deploy` strips `password`, `private_key`, `personal_access_token`, `secret_access_key`, `access_key_id`, `oauth_client_secret`, `oauth_client_id` and `api_key` from `datasources.yml` before it builds the archive, so a secret pasted into the file by mistake is still not deployed. - `.zsql` and `.git` are never part of the archive. ## Supported Adapters `athena`, `bigquery`, `clickhouse`, `databricks`, `druid`, `duckdb`, `mysql`, `postgres`, `redshift`, `snowflake`, `sqlite`, `sqlserver`, `trino`. Each has a page under [Adapters](#adapters) describing the dialect it emits. ## Semantic model and data source scoping Each table in your semantic model belongs to **exactly one datasource**: - Relationships are defined **within** a datasource; cross-datasource joins are not supported. - Universes and query plans are built per datasource, and **every query resolves to exactly one datasource**. The SQL you get back targets that one engine. - Compound measures and automatic data blending operate **inside a single datasource**: you cannot build one measure that mixes fields from Snowflake and Postgres, for example. What multiple datasources are for is **routing**, not blending. Define the same tables and measures in more than one datasource (a Snowflake warehouse of record and a ClickHouse hot tier, say), and the planner picks the datasource that can satisfy the whole query, preferring by partition, tier, and cost. See [Semantic Routing](#semantic-routing). A project can also model unrelated domains in different warehouses, as long as no single query needs both. The response tells you which datasource was chosen, so your application knows which connection to run the SQL on. ## Development engines DuckDB and SQLite are convenient for local development and tests: a file path, no server. The model is the same; only the `adapter` and the dialect of the generated SQL change. SQLite lacks some SQL that analytical queries lean on (window functions on older builds, date truncation), so treat it as a demo target and use a warehouse adapter for the real thing. ## Best Practices 1. **Use descriptive keys**, for example `warehouse`, not `db1`. The key is what table files refer to 2. **Set tiers to match the engine**: hot for the fast tier, cold for object-store engines 3. **Keep secrets out of `datasources.yml`**; nothing in the project needs a working credential 4. **Declare partitions** on tables that hold a slice of the data, so routing has something to work with 5. **Run `zsql check`** after editing; an unsupported adapter or tier fails there, not at deploy ## Next Steps - [Managing Datasources](#datasources), the CLI side - [Adapters](#adapters), what each dialect emits - [Semantic Routing](#semantic-routing), how the planner chooses a datasource - [Cost optimization](#cost-optimization) ### Expressions Learn how to define field expressions using SQL, lookups, and arrays. ## Learning Objectives After completing this guide, you will be able to: - Write SQL expressions for dimensions and measures - Use lookup expressions for dimension fields - Configure primary key expressions - Work with array expressions ## Expression Types ### SQL Expressions SQL expressions are the most common type. They define how to query the field from the database. **For Dimensions:** ```yaml - type: dimension name: Customer Name data_type: string expression: sql: customer_name ``` **For Measures:** ```yaml - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` **Complex SQL:** ```yaml - type: dimension name: Full Name data_type: string expression: sql: CONCAT(first_name, ' ', last_name) ``` **With CASE statements:** ```yaml - type: dimension name: Order Status data_type: string expression: sql: | CASE WHEN status = 'P' THEN 'Pending' WHEN status = 'C' THEN 'Completed' WHEN status = 'X' THEN 'Cancelled' ELSE 'Unknown' END ``` ## Shorthand For simple SQL only, you can use a string instead of a nested mapping (same as `expression: { sql: ... }`): ```yaml - type: dimension name: Order ID data_type: integer expression: order_id - type: measure name: Total Revenue data_type: decimal expression: sum(amount) ``` Use the mapping form when you need `lookup`, `primary_key`, or `array`. ## Primary Keys Mark a dimension as a primary key. This helps with query optimization and ensures uniqueness. ```yaml - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id ``` **When to use:** - Unique identifiers (IDs, codes) - One row per entity - Helps planner optimize joins ## Lookup Expressions Lookup expressions indicate that a dimension field is a lookup/dimension field. This is optional but can help with query planning. ```yaml - type: dimension name: Product Category data_type: string expression: lookup: true sql: category_name ``` **When to use:** - Dimension fields (not measures) - Fields used for grouping/filtering - Optional - SQL expressions work without it ## Array Expressions Array expressions indicate that a field contains array values. ```yaml - type: dimension name: Tags data_type: string expression: array: true sql: tag_array ``` **When to use:** - Database columns that are arrays - JSON array fields - Multi-value dimensions ## Measure Expressions Measures must include aggregation functions: **Sum:** ```yaml - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` **Average:** ```yaml - type: measure name: Average Order Value data_type: decimal expression: sql: avg(amount) ``` **Count:** ```yaml - type: measure name: Order Count data_type: integer expression: sql: count(*) ``` **Count Distinct:** ```yaml - type: measure name: Unique Customers data_type: integer expression: sql: count(distinct customer_id) ``` **Min/Max:** ```yaml - type: measure name: First Order Date data_type: date expression: sql: min(order_date) - type: measure name: Last Order Date data_type: date expression: sql: max(order_date) ``` ## Expression Best Practices 1. **Use column names directly** for simple fields 2. **Include aggregation** for all measures (sum, avg, count, etc.) 3. **Mark primary keys** for unique identifiers 4. **Use SQL functions** for transformations (CONCAT, CASE, etc.) 5. **Keep expressions readable** - complex logic should be in views or calculated columns ## Common Patterns ### Concatenation ```yaml - type: dimension name: Full Address data_type: string expression: sql: CONCAT(street, ', ', city, ', ', state, ' ', zip) ``` ### Date Formatting ```yaml - type: dimension name: Order Month data_type: string expression: sql: TO_CHAR(order_date, 'YYYY-MM') ``` ### Null Handling ```yaml - type: dimension name: Customer Name data_type: string expression: sql: COALESCE(customer_name, 'Unknown') ``` ### Conditional Logic ```yaml - type: dimension name: Customer Segment data_type: string expression: sql: | CASE WHEN total_orders > 100 THEN 'VIP' WHEN total_orders > 50 THEN 'Regular' ELSE 'New' END ``` ## Next Steps - Learn about [relationships](#cardinality) to join tables - Read the [SQL expressions deep-dive](#sql-expressions) - Explore [lookup expressions](#lookup-expressions) - Understand [array expressions](#array-expressions) ### Array Expressions Work with array/multi-value fields. ## Overview Array expressions indicate that a field contains array values. Use for database columns that are arrays or JSON arrays. ## Basic Syntax ```yaml expression: array: true sql: array_column ``` ## When to Use Use array expressions for: - Database array columns (PostgreSQL arrays) - JSON array fields - Multi-value dimensions ## Example ```yaml - type: dimension name: Tags data_type: string expression: array: true sql: tag_array ``` ## PostgreSQL Arrays For PostgreSQL array columns: ```yaml - type: dimension name: Product Tags data_type: string expression: array: true sql: tags # PostgreSQL array column ``` ## JSON Arrays For JSON array fields: ```yaml - type: dimension name: Categories data_type: string expression: array: true sql: json_extract_path_text(categories, 'items') # JSON array ``` ## Best Practices 1. **Use for multi-value fields** - When a field can have multiple values 2. **Test array handling** - Verify arrays work in your queries 3. **Document array structure** - Add descriptions explaining array format ## Next Steps - Learn about [SQL expressions](#sql-expressions) - Explore [lookup expressions](#lookup-expressions) - Understand [dimensions](#dimensions) ### Lookup Expressions Indicate dimension fields with lookup expressions. ## Overview Lookup expressions are optional indicators that a dimension field is a lookup/dimension field. They help the planner understand field semantics. ## Basic Syntax ```yaml expression: lookup: true sql: column_name ``` ## When to Use Lookup expressions are optional but can help with: - Query planning optimization - Field semantics documentation - Planner understanding of field types ## Example ```yaml - type: dimension name: Product Category data_type: string expression: lookup: true sql: category_name ``` ## Notes - **Optional** - SQL expressions work without `lookup: true` - **Dimensions only** - Not used for measures - **Semantic indicator** - Helps planner understand field purpose ## Complete Example ```yaml fields: - type: dimension name: Customer Name data_type: string expression: lookup: true sql: customer_name - type: dimension name: Product Category data_type: string expression: lookup: true sql: category_name - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) # No lookup for measures ``` ## Best Practices 1. **Use for clarity** - Makes field semantics explicit 2. **Optional** - Not required, but can help documentation 3. **Dimensions only** - Don't use for measures ## Next Steps - Learn about [SQL expressions](#sql-expressions) - Explore [array expressions](#array-expressions) - Understand [dimensions](#dimensions) ### SQL Expressions Define field expressions using SQL. ## Overview SQL expressions are the most common way to define how fields query the database. They allow you to write SQL directly or reference database columns. ## Basic Syntax ```yaml expression: sql: column_name ``` **Shorthand:** `expression: column_name` is equivalent to `expression: { sql: column_name }`. Use the mapping form when you need `lookup`, `primary_key`, or `array`. ## Simple Column Reference Reference a database column directly: ```yaml - type: dimension name: Customer Name data_type: string expression: sql: customer_name ``` ## SQL Functions Use SQL functions for transformations: ```yaml - type: dimension name: Full Name data_type: string expression: sql: CONCAT(first_name, ' ', last_name) ``` ## CASE Statements Use CASE for conditional logic: ```yaml - type: dimension name: Order Status data_type: string expression: sql: | CASE WHEN status = 'P' THEN 'Pending' WHEN status = 'C' THEN 'Completed' WHEN status = 'X' THEN 'Cancelled' ELSE 'Unknown' END ``` ## Aggregation (Measures Only) Measures must include aggregation functions: ```yaml - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` **Common aggregations:** - `sum()` - Sum of values - `avg()` - Average of values - `count()` - Count of rows - `count(distinct column)` - Count of distinct values - `min()` - Minimum value - `max()` - Maximum value ## Multi-line SQL Use YAML literal block for multi-line SQL: ```yaml - type: dimension name: Complex Calculation data_type: decimal expression: sql: | CASE WHEN amount > 1000 THEN 'High Value' WHEN amount > 500 THEN 'Medium Value' ELSE 'Low Value' END ``` ## Database-Specific Functions Use adapter-specific SQL functions: ```yaml # PostgreSQL - type: dimension name: Order Month data_type: string expression: sql: TO_CHAR(order_date, 'YYYY-MM') # Snowflake - type: dimension name: Order Month data_type: string expression: sql: TO_VARCHAR(order_date, 'YYYY-MM') ``` ## Null Handling Handle NULL values: ```yaml - type: dimension name: Customer Name data_type: string expression: sql: COALESCE(customer_name, 'Unknown') ``` ## Date Functions Transform dates: ```yaml - type: dimension name: Order Year data_type: integer expression: sql: EXTRACT(YEAR FROM order_date) ``` ## Best Practices 1. **Keep it simple** - Prefer simple column references 2. **Use database functions** - Leverage SQL capabilities 3. **Handle NULLs** - Use COALESCE or CASE for NULL handling 4. **Test expressions** - Verify SQL works in your database 5. **Document complex logic** - Add descriptions for complex expressions ## Common Patterns ### Concatenation ```yaml - type: dimension name: Full Address data_type: string expression: sql: CONCAT(street, ', ', city, ', ', state, ' ', zip) ``` ### Date Formatting ```yaml - type: dimension name: Order Month data_type: string expression: sql: TO_CHAR(order_date, 'YYYY-MM') ``` ### Conditional Logic ```yaml - type: dimension name: Customer Segment data_type: string expression: sql: | CASE WHEN total_orders > 100 THEN 'VIP' WHEN total_orders > 50 THEN 'Regular' ELSE 'New' END ``` ### Aggregation with Filter ```yaml - type: measure name: Completed Orders data_type: integer expression: sql: count(CASE WHEN status = 'completed' THEN 1 END) ``` ## Next Steps - Learn about [lookup expressions](#lookup-expressions) - Explore [array expressions](#array-expressions) - Understand [primary keys](#primary-keys) ### Extended Blending Groups Use extended blending groups to treat the same logical dimension across multiple tables as one for querying. ## Overview When the same business dimension (e.g. Order Date, Ship Date, Activity Date) exists in several tables with different column names or grains, you can assign a shared **`extended_blend_group`** on each. The planner can then blend these dimensions so that a single query can group or filter by that concept across tables. ## When to Use Extended Blending - **Cross-table date alignment**: e.g. “Order Date” in `orders` and “Order Date” in `shipments` or `returns` so users can filter by one order-date concept. - **Shared attributes in a star or snowflake**: The same dimension name and meaning in multiple fact tables; you want one blended dimension instead of one per table. - **Multi-fact queries**: specs that slice by a common dimension (date, customer, product) across several fact tables. ## Setting Up `extended_blend_group` Use the same `extended_blend_group` value on each dimension that should blend: ```yaml # tbl.orders.yml fields: - type: dimension name: Order Date data_type: date extended_blend_group: order_date_blend grains: - day - week - month - quarter - year expression: sql: order_date ``` ```yaml # tbl.shipments.yml fields: - type: dimension name: Order Date data_type: date extended_blend_group: order_date_blend grains: - day - week - month - quarter - year expression: sql: ship_date ``` The string is case-insensitive. Only dimensions in the same group blend together. ## Example: Multiple Fact Tables Sales, returns, and inventory snapshots each have a date. Define one blend group for “activity date”: ```yaml # tbl.sales.yml fields: - type: dimension name: Activity Date data_type: date extended_blend_group: activity_date grains: - day - week - month - quarter - year expression: sql: sale_date # tbl.returns.yml fields: - type: dimension name: Activity Date data_type: date extended_blend_group: activity_date grains: - day - week - month - quarter - year expression: sql: return_date # tbl.inventory.yml fields: - type: dimension name: Activity Date data_type: date extended_blend_group: activity_date grains: - day - week - month - quarter - year expression: sql: snapshot_date ``` Queries that use “Activity Date” can then apply it across sales, returns, and inventory according to the relationships and paths in the model. ## Requirements and Best Practices 1. **Same `data_type`**: All dimensions in a group should have the same type (e.g. all `date` or all `date_time`). 2. **Compatible grains**: For `date`/`date_time`, use grains that make sense for all members (e.g. day, week, month, quarter, year). 3. **Same universe/branch**: Blending applies within a branch; dimensions in different branches do not blend. 4. **Meaningful group names**: Use a group name that reflects the business concept (e.g. `order_date_blend`, `activity_date`). 5. **Avoid over-blending**: Only put dimensions in a group when they truly represent the same concept; mixing different semantics can produce confusing results. ## How automatic data blending works (developer view) Extended blend groups are one part of 0sql's **automatic data blending** system. In industry terms, this is sometimes called **data blending** or **multipass SQL**; 0sql uses the terms *automatic data blending* and *extended blend groups*. At a high level: - There is a single **logical key** for a concept like “activity date” in the semantic layer. - Each fact table exposes one or more physical columns that implement that key, and you tie those columns together using a shared `extended_blend_group`. - At query time, the planner knows which physical column to use for each fact, and can safely aggregate and **merge** results on the logical key. When a query asks for measures from multiple facts that share a blend group: 1. **Universe selection** picks the relevant universes/facts. 2. For each fact/universe, the planner builds a node in a query graph that aggregates that fact **to the lowest common dimensionality** (for example, by the blended date and any other shared dimensions). 3. The planner then adds a final **merge node** that joins these aggregate nodes on the blended dimensions. The statement uses Common Table Expressions (CTEs) to represent this graph. The sketch below is illustrative, with the names simplified; a real blend, verbatim from the engine, is on [Cross-fact blend](#cross-fact-blend): ```sql WITH store_sales_agg AS ( SELECT activity_date, SUM(store_amount) AS store_revenue FROM store_sales GROUP BY activity_date ), web_sales_agg AS ( SELECT activity_date, SUM(web_amount) AS web_revenue FROM web_sales GROUP BY activity_date ), final AS ( SELECT COALESCE(s.activity_date, w.activity_date) AS activity_date, store_revenue, web_revenue FROM store_sales_agg s FULL OUTER JOIN web_sales_agg w ON s.activity_date = w.activity_date ) SELECT * FROM final; ``` You do not need to write this SQL yourself, the semantic model (relationships, grains, and `extended_blend_group` values) gives the planner enough information to construct it automatically. ## Relation to Relationships and Paths Blending does not replace [relationships](#cardinality). Tables still need proper `one_to_many`, `many_to_one`, or `one_to_one` relationships. Extended blend groups tell the planner how to treat dimensions that live in different tables as one for grouping and filtering when paths exist. ## Next Steps - [Dimensions](#dimensions) - [Relationships](#cardinality) ### Fields and Types Understand dimensions, measures, data types and field metadata in 0sql. ## Learning Objectives After completing this guide, you will be able to: - Distinguish between dimensions and measures - Choose appropriate data types - Configure field metadata (format, display types, synonyms, hidden) - Understand when to use different field options ## Dimensions vs Measures ### Dimensions Dimensions are categorical fields used for grouping and filtering. They represent attributes of your data. **Characteristics:** - Used in GROUP BY clauses - Can be filtered - Typically lower cardinality - Examples: Customer Name, Product Category, Order Date **Example:** ```yaml - type: dimension name: Customer Name data_type: string expression: sql: customer_name ``` ### Measures Measures are aggregatable quantities. They represent quantitative values that can be summed, averaged, or counted. **Characteristics:** - Used with aggregation functions (SUM, AVG, COUNT) - Cannot be used in GROUP BY - Examples: Total Revenue, Order Count, Average Order Value **Example:** ```yaml - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` ## Field Properties Beyond `type`, `name`, `data_type` and `expression`, a field carries metadata. 0sql validates these keys at deploy and stores them on the field. None of them changes the SQL. Some of them come back from `GET .../fields`: `description`, `synonyms`, `tags` and `hidden` are in that response; `format` and `display_type` are not, so your client reads those from the model it deployed. See [Discovery](#discovery-explore-fields-and-tables). ### Display Types `display_type` tells your client what kind of value this is: - **default**: plain value - **html**: an HTML fragment - **url**: a link - **email**: an email address - **phone_number**: a phone number - **image**: an image URL ```yaml - type: dimension name: Product Image data_type: string display_type: image expression: sql: image_url ``` ### Format `format` describes how your client should print the value, as a shortcut string or a mapping. It rides along with the field's metadata; nothing in SQL generation uses it. See [Field format](#field-format). ```yaml - type: measure name: Total Revenue data_type: decimal format: currency:2 expression: sql: sum(amount) ``` ### Hidden Fields `hidden: true` hides a field from discovery: the fields endpoint and `zsql fields` leave it out unless asked for hidden fields, and `explore` never suggests it. A spec can still reference it by name, and compound measures and calculations can build on it: ```yaml - type: dimension name: Internal ID data_type: integer hidden: true expression: sql: internal_id ``` ### Synonyms Alternative names for a field. A field reference in a spec resolves by uid, name or synonym, and `GET .../fields?q=` matches synonyms too, so a caller that says `client name` lands on `Customer Name`. ```yaml - type: dimension name: Customer Name data_type: string synonyms: - client name - buyer name - account name expression: sql: customer_name ``` ### Grains (Date/DateTime Only) Specify supported granularities for date fields: ```yaml - type: dimension name: Order Date data_type: date grains: - day - week - month - quarter - year expression: sql: order_date ``` Available grains: `raw`, `millisecond`, `second`, `minute`, `hour`, `day`, `week`, `month`, `quarter`, `year`. The list advertises what the field supports; the grain itself is applied at request time with a `truncate` decorator (see [Projections](#projections-and-decorators)). ### Tags `tags` are free-form labels. [Security policies](#row-level-security) trigger on them, so a `pii` tag on a field is what makes a policy fire when that field is projected. ## Data Types Every field declares a `data_type` that matches the warehouse column (or, for a measure, the aggregate's result): | `data_type` | Warehouse types | Use for | |---|---|---| | `string` | VARCHAR, TEXT | names, codes, categories, status values | | `integer` | INT | counts, quantities, small IDs | | `bigint` | BIGINT | large IDs, epoch timestamps | | `decimal` | DECIMAL, NUMERIC, FLOAT, DOUBLE | money, ratios, averages, measurements | | `date` | DATE | dates without time; takes `grains` | | `date_time` | TIMESTAMP, DATETIME | timestamps; takes `grains` | | `boolean` | BOOLEAN | flags | | `binary` | BLOB, BINARY | rare; prefer a `string` URL or path | Date types also declare which `grains` callers may group by (see [Grains](#grains-datedatetime-only) above): ```yaml - type: dimension name: Order Date data_type: date grains: [day, week, month, quarter, year] expression: sql: order_date - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` Full reference: [Data Types](#data-types). ## Complete Example ```yaml fields: # Primary key dimension - type: dimension name: Order ID description: Unique order identifier data_type: integer expression: primary_key: true sql: order_id # Date dimension with multiple grains - type: dimension name: Order Date description: Date when order was placed data_type: date grains: - day - week - month - quarter - year expression: sql: order_date # High-cardinality dimension - type: dimension name: Customer ID description: Customer identifier data_type: integer expression: sql: customer_id # Formatted measure with synonyms - type: measure name: Total Revenue description: Sum of all order amounts data_type: decimal format: currency:2 synonyms: - total sales - gross revenue - sales amount expression: sql: sum(amount) # Count measure - type: measure name: Order Count description: Number of orders data_type: integer expression: sql: count(*) ``` ## Best Practices 1. **Use appropriate data types** - Match your database column types 2. **Add descriptions** - Help users understand what each field represents 3. **Set grains for dates** - Enable flexible date grouping 4. **Set format** on money and ratio measures so your client prints them consistently 5. **Mark primary keys** - Helps with query optimization ## Next Steps - Learn about [expressions](#expressions) (SQL, lookups, arrays) - Read the [dimensions deep-dive](#dimensions) - Read the [measures deep-dive](#measures) - Explore [data types reference](#data-types) ### Data Types The eight data types a field can declare, and which warehouse types they map to. ## Reference | `data_type` | Warehouse types | Use for | Notes | |---|---|---|---| | `string` | VARCHAR, TEXT, CHAR | names, descriptions, codes, categories, status values | Keep order numbers and SKUs as strings unless you do math on them | | `integer` | INT, INTEGER, SMALLINT | counts, quantities, years, small IDs | 32-bit range | | `bigint` | BIGINT | large IDs, epoch timestamps, large counts | 64-bit range | | `decimal` | DECIMAL, NUMERIC, FLOAT, DOUBLE | money, percentages, ratios, averages, measurements | Precision follows the warehouse; use it for every monetary value | | `date` | DATE | dates without a time part | Declares `grains`: `day`, `week`, `month`, `quarter`, `year` | | `date_time` | TIMESTAMP, DATETIME | timestamps, event times | Declares `grains`, plus `raw`, `second`, `minute`, `hour` | | `boolean` | BOOLEAN, BOOL | flags, yes/no fields | Values `true`, `false`, `NULL` | | `binary` | BLOB, BINARY, BYTEA | binary payloads | Rare; store a URL or path as a `string` instead | A measure's `data_type` describes the aggregate's result: `count(*)` is an `integer`, `sum(amount)` a `decimal`. ## Examples ```yaml - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id - type: dimension name: Order Date data_type: date grains: [day, week, month, quarter, year] expression: sql: order_date - type: dimension name: Created At data_type: date_time grains: [raw, hour, day, week, month] expression: sql: created_at - type: measure name: Total Revenue data_type: decimal format: currency:2 expression: sql: sum(amount) ``` ## Grains `date` and `date_time` fields list the granularities callers can group and filter by. When a spec asks for a grain with a `truncate` decorator, 0sql truncates the column to it in the generated SQL; the list advertises which grains make sense for the field. | Grain | Applies to | |---|---| | `raw` | `date_time` only: the untruncated value | | `millisecond`, `second`, `minute`, `hour` | `date_time` only | | `day`, `week`, `month`, `quarter`, `year` | both | ## Guidelines - **Match the column.** Mismatches surface as cast errors when your application runs the SQL, not at deploy. - **`decimal` for money, never a float column type name.** Currency formatting is a `format` hint for your client, not the type. See [Field format](#field-format). - **List only useful grains.** A daily inventory table has no business offering `hour`. - **Prefer `integer` for IDs** unless the column is a BIGINT. ## Next Steps - [Fields and Types](#fields-and-types) - properties every field can set - [Dimensions](#dimensions) and [Measures](#measures) - [Field format](#field-format) - format metadata ### Date and DateTime Dimensions Use `date` and `date_time` dimensions for time-based grouping and filtering. ## Overview - **`date`**: Date without time (e.g. order date, birth date). Format: YYYY-MM-DD. - **`date_time`**: Timestamp with time (e.g. created_at, event time). Format: YYYY-MM-DD HH:MM:SS. Define them as dimensions and set **grains** to advertise how callers can group and filter. ## Grains Grains are the granularities you allow for a date or date_time dimension. ```yaml - type: dimension name: Order Date data_type: date grains: - day - week - month - quarter - year expression: sql: order_date ``` **Available grains:** | Grain | Typical use | `date` | `date_time` | |-----------|------------------|--------|-------------| | `raw` | Original value | n/a | ✓ | | `second` | Second-level | n/a | ✓ | | `minute` | Minute-level | n/a | ✓ | | `hour` | Hour-level | n/a | ✓ | | `day` | Day-level | ✓ | ✓ | | `week` | Week-level | ✓ | ✓ | | `month` | Month-level | ✓ | ✓ | | `quarter` | Quarter-level | ✓ | ✓ | | `year` | Year-level | ✓ | ✓ | For `date_time`, you can include `raw`, `second`, `minute`, and `hour` when you need sub-day precision. ## Date Dimension Example ```yaml - type: dimension name: Order Date description: Date when order was placed data_type: date grains: - day - week - month - quarter - year expression: sql: order_date ``` ## DateTime Dimension Example ```yaml - type: dimension name: Created At description: When the record was created data_type: date_time grains: - raw - hour - day - week - month expression: sql: created_at ``` ## When to Use date vs date_time - **`date`**: Order date, ship date, birth date, any dimension where time-of-day does not matter. - **`date_time`**: Created at, updated at, event time, any dimension where hour/minute/second can be used for grouping or filtering. ## Grouping and Formatting - A grain is applied at request time with a `truncate` decorator on the projection; the SQL you get back carries `DATE_TRUNC('month', ...)` in the datasource's dialect. See [Projections](#projections-and-decorators). - Set a [field format](#field-format) such as `date:short` when your client should print the value a particular way; it is metadata, not SQL. ## Best Practices 1. **Set grains**: only include grains you need; they tell callers what grouping makes sense. 2. **Prefer `date` when time is irrelevant**: Simpler and more efficient. 3. **Use `date_time` for events**: when callers need hour or minute level. ## Next Steps - [Dimensions](#dimensions) - [Data types](#data-types) - [Field format](#field-format) ### Field format `format` describes how a field's values should be presented. 0sql validates it at deploy and stores it on the field. Nothing in SQL generation reads it: the statement you get back returns raw values, and your client applies the format. It is not one of the keys `GET .../fields` returns, so your client gets it from the model you deployed, not from the discovery endpoints. ## Overview Set **`format`** on a dimension or measure in table YAML, in one of two equivalent forms: 1. **Shortcut string**, compact colon-separated syntax such as `currency:2` or `percent:1` 2. **Mapping**, an explicit `type` plus options Both normalize to the same stored JSON, for example `{"type":"currency","precision":2}`. Use the shortcut for the common cases and the mapping when you need several options or want the review diff to be obvious. ## Shortcut syntax `type[:precision[:unit]][:abbreviate]` | Shortcut | Stored format | |----------|---------------| | `number:2` | `{ type: number, precision: 2 }` | | `number:2:abbreviate` | `{ type: number, precision: 2, abbreviate: true }` | | `currency:2` | `{ type: currency, precision: 2 }` | | `currency:2:USD` | `{ type: currency, precision: 2, unit: "USD" }` | | `currency:2:abbreviate` | `{ type: currency, precision: 2, abbreviate: true }` | | `percent:1` | `{ type: percent, precision: 1 }` | | `date:short` | `{ type: date, pattern: "%b %d, %Y" }` | | `date:long` | `{ type: date, pattern: "%B %d, %Y" }` | | `date:iso` | `{ type: date, pattern: "%Y-%m-%d" }` | | `date:%m/%d/%Y` | `{ type: date, pattern: "%m/%d/%Y" }`, any strftime pattern | | `datetime:short` | `{ type: datetime, pattern: "%b %d, %Y %I:%M %p" }` | | `datetime:iso` | `{ type: datetime, pattern: "%Y-%m-%d %H:%M:%S" }` | | `html:{{value}}` | `{ type: html, template: "..." }` | | `javascript:formatValue` | `{ type: javascript, function: "formatValue" }` | Types: `number`, `currency`, `percent`, `date`, `datetime`, `html`, `javascript`. A shortcut with an unknown type is dropped silently and the field has no format. If a shortcut is ambiguous in your YAML editor, quote it: `format: "percent:2"`. `html` and `javascript` are stored like any other type. What they mean is up to your rendering layer; 0sql does not evaluate templates or functions. ## Mapping syntax | type | Keys | |------|------| | `number` | `precision`, `abbreviate` | | `currency` | `precision`, `unit`, `abbreviate` | | `percent` | `precision` | | `date`, `datetime` | `pattern` (strftime) | | `html` | `template` (`{{value}}` stands for the value) | | `javascript` | `function` | Normalization: blank values are dropped, `precision` is coerced to an integer, and `abbreviate` is kept only when true. Other keys are stored as given. ## Examples ```yaml - type: measure name: Store Net Paid data_type: decimal format: currency:2 expression: sql: sum(ss_net_paid) - type: measure name: Profit Margin data_type: decimal format: type: percent precision: 2 expression: sql: sum(ss_net_profit) / nullif(sum(ss_net_paid), 0) - type: measure name: Store Quantity data_type: integer format: number:0 expression: sql: sum(ss_quantity) - type: dimension name: Date data_type: date format: date:short expression: sql: d_date ``` A client that reads the stored format would show these as `$1,234.56`, `45.67%` (for a stored ratio of `0.4567`), `1,234` and `Jan 05, 2026`. ## Format vs display type `display_type` says what kind of thing the value is (`url`, `email`, `image`, `html`, `phone_number`); `format` says how to print it (numbers, dates, templates). Both are metadata, both can sit on the same field, and neither changes the SQL. ## Validation `zsql check` and `zsql deploy` reject a `format` that is neither a string nor a mapping with `Field errors: format must be a shortcut string or a mapping`. There is no default format: a field without `format` carries none, and your client decides how to print a `decimal` or a `date`. ## Best practices 1. Use **`currency:2`**, or a mapping with `unit`, for monetary measures 2. Use **`percent`** for ratios stored as decimals (0 to 1) 3. Use **`number:0`** for integer counts 4. Keep business logic in SQL; `format` is presentation only ## Next Steps - [Dimensions](#dimensions) and [measures](#measures) - [Data types](#data-types) - [Display types](#display-types) - [Field discovery](#discovery-explore-fields-and-tables), for the field keys the API does return ### Compound Measures Measures written as a formula over other measures and dimensions by name, across fact domains. ## Overview A compound measure's expression references other fields in square brackets instead of table columns: ```yaml - type: measure name: Profit Margin data_type: decimal format: percent:1 expression: sql: ([Total Revenue] - [Total Cost]) / nullif([Total Revenue], 0) ``` The formula is resolved at **query time**. Each referenced measure is planned as itself, on whatever table serves it best for the dimensions requested, and the formula is applied to the results. That is what makes a compound measure different from a [standard measure](#measures) with arithmetic inside one aggregate: the pieces don't have to come from the same table, or the same fact. ## Spanning fact domains Referenced measures can live on **different fact tables in different domains**. `[Store Revenue]` from store sales, `[Web Revenue]` from web sales, `[Support Tickets]` from a service fact: a compound measure can combine any of them. ```yaml - type: measure name: Tickets per 1000 Orders data_type: decimal expression: sql: 1000.0 * [Support Tickets] / nullif([Order Count], 0) ``` When a query groups this by `Customer Segment` and `Month`, the planner aggregates orders in the orders universe, tickets in the support universe, each to that grain, blends the two on the conformed dimensions with a full outer join, and then evaluates the division. No row-level join between the facts is ever generated, so the result cannot fan out. This is the same automatic blending every multi-measure query gets; the compound measure just packages it under one name. ## Referencing dimensions A formula may also reference **dimensions**, most often inside a conditional aggregate or a `CASE`: ```yaml - type: measure name: Premium Share of Revenue data_type: decimal format: percent:1 expression: sql: sum(case when [Customer Tier] = 'Premium' then [Order Amount] else 0 end) / nullif([Total Revenue], 0) ``` ```yaml - type: measure name: Weekend Orders data_type: integer expression: sql: count(case when [Day of Week] in ('Sat', 'Sun') then 1 end) ``` Here `[Customer Tier]` and `[Day of Week]` are dimensions. A referenced dimension is resolved separately within each universe the formula's measures are defined on, so it doesn't have to sit on the fact table itself: it can be anywhere along the join path. `[Customer Tier]` reached from store sales through Customer, and from web sales through the same Customer dimension, resolves to the right joined column in each sub-query. ## The restrictions There are two, and neither is about tables or universes. **Every field in the formula must belong to the same datasource.** The formula is evaluated by one engine, so `[Store Revenue]` on Snowflake and `[Support Tickets]` on Postgres cannot be combined in a compound measure. Different fact domains within one datasource are fine, which is the common case. **Every dimension a formula references must be reachable from every fact domain the formula involves.** Reachable means anywhere in that fact's universe, along its join path, not necessarily on the fact table. A compound measure that reads `[Store Revenue]` and `[Web Revenue]` may reference `[Customer Tier]` if both the store sales universe and the web sales universe can reach the Customer dimension, however many joins away. A dimension that only one side can reach makes the formula unanswerable, and 0sql refuses it at validation rather than guessing. This rule is about dimensions written *into the formula*. The dimensions a user *groups by* at query time are handled differently: a component that can't reach one of them is auto-leveled, as described next. There is no requirement that the referenced measures share a table or a universe. ## Auto-leveled compound measures When a compound measure spans fact domains, the query's dimensions may not all be reachable from every component. Rather than refuse the query, 0sql **auto-levels** the components that can't reach a dimension: that dimension is excluded from that component's aggregation, and the final blend joins only on the dimensions the components have in common. Take `Tickets per 1000 Orders` from above, and a query grouped by `Month` and `Contact Channel`: - `[Support Tickets]` lives on the support fact, which reaches `Contact Channel`. It aggregates by `Month` and `Contact Channel`. - `[Order Count]` lives on the orders fact, which has no path to `Contact Channel`. It is auto-leveled: `Contact Channel` is dropped from its grouping and it aggregates by `Month` only. - The two are joined on the common dimension, `Month`, and the formula is applied. Each channel's ticket count is divided by that month's total order count. | Month | Contact Channel | Support Tickets | Order Count (leveled) | Tickets per 1000 Orders | |---|---|---|---|---| | 2026-08 | Chat | 420 | 91,000 | 4.6 | | 2026-08 | Email | 310 | 91,000 | 3.4 | | 2026-08 | Phone | 180 | 91,000 | 2.0 | The leveled component repeats across the dimension it can't see, which is exactly the fixed-denominator behavior a rate like this wants. It is the same mechanism as an [exclusion](#exclusions), applied automatically and only to the components that need it. A component that can reach every dimension is never leveled. Two things to keep in mind: - **Auto-leveling applies to query-time grouping, not to dimensions written into the formula.** A dimension referenced inside the expression must be reachable from every component (the rule above). - **A leveled component is a total over the missing dimension, not a per-row value.** For a rate that is the intent. For a sum or difference across domains, groups are only meaningful on dimensions both sides reach, so be explicit in the measure's description about which dimensions it should be grouped by. ## Chaining Compound measures can reference other compound measures. The dependency chain is resolved when the query is built, so definition order in YAML doesn't matter. Circular references (`A` → `B` → `A`) are invalid. ```yaml fields: - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) - type: measure name: Total Cost data_type: decimal expression: sql: sum(cost) - type: measure name: Profit data_type: decimal expression: sql: "[Total Revenue] - [Total Cost]" - type: measure name: Profit Margin data_type: decimal format: percent:1 expression: sql: "[Profit] / nullif([Total Revenue], 0)" ``` Bracketed references mix freely with SQL functions and constants: ```yaml - type: measure name: Adjusted Profit Margin data_type: decimal format: percent:1 expression: sql: ([Profit] - coalesce([Refunds], 0)) / nullif([Total Revenue], 0) ``` ## Where the formula lives Compound measures are the modeled, governed form of a formula: deployed with the project and available to every caller by name. The same `[Name]@m` / `[Name]@d` syntax is available at request time as a [calculation](#calculations) in the spec, where the `@m` and `@d` suffixes disambiguate a name shared by a dimension and a measure. A calculation your application keeps sending is a candidate to promote here. [Decorators](#projections-and-decorators) (year-over-year, moving average, percent of total) work on compound measures as they do on any other measure. ## Guidelines 1. **Reference by name, not by table.** The planner picks the table; the formula shouldn't assume one. 2. **Guard denominators** with `nullif(..., 0)`. 3. **Stay within one datasource, and keep dimension references to ones every involved fact can reach.** If a dimension is specific to one fact, model the conditional part as a standard measure on that fact and reference that measure instead. 4. **Prefer a chain of small measures** over one long formula; each piece is reusable and easier to test. ## Next Steps - [Measures](#measures) - [Calculations](#calculations), the request-time form - [Universe formation](#universe-formation) - [Expressions](#sql-expressions) ### Snapshot Measures Capture point-in-time values for inventory and balance analysis. ## Overview Snapshot measures capture values at specific points in time (beginning or ending of a period). They're useful for **stateful** quantities like inventory, account balances, or membership counts, cases where summing across days is not meaningful, and you care about the value **at the boundary of a period**. ## Requirements 1. **Table must set `snapshot`** – Names the date dimension that drives snapshot behavior. 2. **Measure must specify `snapshot`** – `ending` or `beginning`. The table's `snapshot` dimension can: - Reference a date dimension defined **on the same table**, or - Refer to a date dimension in the table’s **universe**, as long as the association is unambiguous (for example, a single “As Of Date” dimension that always joins to this fact). ## Table Configuration Set the snapshot date dimension on the table: ```yaml name: Inventory physical_name: inventory datasource: warehouse cost: 100 snapshot: Date # Required for snapshot measures: the date dimension the table is snapped by fields: - type: dimension name: Date data_type: date expression: sql: snapshot_date ``` ## Measure Configuration ### Ending Snapshot Capture value at the end of the period: ```yaml - type: measure name: Ending Inventory data_type: integer snapshot: ending expression: sql: sum(quantity) ``` ### Beginning Snapshot Capture value at the beginning of the period: ```yaml - type: measure name: Beginning Inventory data_type: integer snapshot: beginning expression: sql: sum(quantity) ``` ## How It Works When a spec projects a snapshot measure: 1. The planner identifies the **snapshot date dimension** for the table. 2. It filters rows to the requested time period (for example, a specific month or week). 3. It chooses the appropriate snapshot within each bucket: - `snapshot: ending` → the **last** snapshot in the bucket. - `snapshot: beginning` → the **first** snapshot in the bucket. 4. It generates SQL that aggregates the underlying values at those boundary timestamps only, rather than summing every row in the period. This means that: - Grouping by a period (day/week/month/etc.) yields “as of end of period” or “as of start of period” values. - Grouping by non-time dimensions (for example, `Product` or `Store`) still respects the snapshot semantics as long as the query includes a time filter or grain. ## Use Cases ### Inventory Analysis ```yaml name: Inventory snapshot: Date fields: - type: measure name: Ending Inventory snapshot: ending expression: sql: sum(quantity) ``` ### Account Balances ```yaml name: Accounts snapshot: Date fields: - type: measure name: Ending Balance snapshot: ending expression: sql: sum(balance) ``` ## Complete Example ```yaml name: Inventory physical_name: inventory datasource: warehouse cost: 100 snapshot: Date fields: - type: dimension name: Date data_type: date expression: sql: snapshot_date - type: dimension name: Product ID data_type: integer expression: sql: product_id - type: measure name: Beginning Inventory data_type: integer snapshot: beginning expression: sql: sum(quantity) - type: measure name: Ending Inventory data_type: integer snapshot: ending expression: sql: sum(quantity) ``` ## Best Practices 1. **Set `snapshot` on the table** – Required; keep the association with the driving date dimension clear. 2. **Use descriptive names** – e.g. "Ending Inventory" not just "Inventory". 3. **Document purpose** – Note that the measure is point-in-time and should not be summed across long ranges. 4. **Test the SQL** – Add a `tests/*.yml` projecting the snapshot measure by month and assert the boundary-date subquery is there; then compare daily snapshots against month-end values in your warehouse. ## Next Steps - [Measures](#measures) - [Exclusions](#exclusions) - [Inclusions](#inclusions) - [Temporal decorators](#projections-and-decorators) ### Imports Reuse field definitions across tables. ## Overview Imports allow you to define fields once and reuse them across multiple tables. This reduces duplication and ensures consistency. ## How Imports Work 1. **Imports processed first** - Fields from imported files are loaded 2. **Local fields merged** - Local field definitions override or extend imports 3. **Circular dependency detection** - Validated to prevent infinite loops ## Basic Syntax ```yaml imports: - path/to/shared_fields.yml - ../common/dates.yml ``` ## Example ### Shared Fields File **models/common/dates.yml:** ```yaml fields: - type: dimension name: Order Date data_type: date grains: - day - week - month - quarter - year expression: sql: order_date ``` ### Importing Table **models/sales/tbl.orders.yml:** ```yaml name: Orders physical_name: orders datasource: warehouse cost: 100 imports: - ../common/dates.yml fields: - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` ## Overriding Imported Fields Local fields can override imported fields: ```yaml imports: - ../common/dates.yml fields: # Override imported Order Date with custom expression - type: dimension name: Order Date data_type: date expression: sql: custom_order_date # Override imported field ``` ## Extending Imported Fields Add additional fields alongside imports: ```yaml imports: - ../common/dates.yml fields: # Imported fields are included # Plus these additional fields: - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` ## Path Resolution Paths are relative to the importing file: ```yaml # From models/sales/tbl.orders.yml imports: - ../common/dates.yml # models/common/dates.yml - ../inventory/products.yml # models/inventory/products.yml - common/geo_dimensions.yml # models/sales/common/geo_dimensions.yml ``` ## Common Patterns ### Shared Date Dimensions **models/common/dates.yml:** ```yaml fields: - type: dimension name: Order Date data_type: date grains: - day - week - month - quarter - year expression: sql: order_date ``` **Multiple tables import:** ```yaml # models/sales/tbl.orders.yml imports: - ../common/dates.yml # models/inventory/tbl.stock.yml imports: - ../common/dates.yml ``` ### Shared Geography Fields **models/common/geo_dimensions.yml:** ```yaml fields: - type: dimension name: Region data_type: string expression: sql: region_name - type: dimension name: Country data_type: string expression: sql: country_name ``` ## Best Practices 1. **Group related fields** - Create logical groupings (dates, geography, etc.) 2. **Use descriptive paths** - Clear file names and organization 3. **Document imports** - Add comments explaining what's imported 4. **Avoid deep nesting** - Keep import paths simple 5. **Test imports** - Verify imported fields work correctly ## Validation The CLI validates: - **File existence** - Imported files must exist - **Circular dependencies** - Prevents infinite loops - **Field conflicts** - A local field with the same name overrides the imported one ## Troubleshooting ### Import Not Found **Error:** `Import not found: ../common/dates.yml` **Solution:** - Check file path is correct (relative to importing file) - Verify file exists - Check file permissions ### Circular Dependency **Error:** `Circular import detected: a.yml -> b.yml -> a.yml` **Solution:** - Review import chain - Remove circular references - Restructure shared fields ## Next Steps - Learn about [tables](#tables) - Read about [fields](#dimensions) - Explore [table schema reference](#tables) ### Partitions Define data availability constraints for query routing. ## Overview Partitions tell the planner about data availability constraints. This helps route queries to the correct tables when multiple tables can answer the same question. ## Supported Predicates The predicates that rank a table are `between`, `greater_than`, `greater_than_or_equal_to`, `less_than` and `less_than_or_equal_to` (number or date dimensions) and `in_list` (any dimension). Any other predicate is accepted with a warning and does not influence routing. Write them under `partitions:` as a list; the singular `partition:` key is also accepted. A partition takes `dimension`, `predicate`, `filter_value` and, for `between`, `filter_value_end`. A `description` is accepted and ignored by the loader; it documents the constraint for readers. ### between Date range partition. ```yaml partitions: - dimension: Order Date predicate: between filter_value: 24m # 24 months ago filter_value_end: 1d # yesterday description: "Rolling 2 year window" ``` ### in_list List of values partition. ```yaml partitions: - dimension: Region predicate: in_list filter_value: us-east, us-west, europe description: "Table only has 3 regions of 5" ``` ## Date Range Partitions Most common use case: rolling window tables. ```yaml name: Recent Sales physical_name: recent_sales datasource: warehouse cost: 10 partitions: - dimension: Order Date predicate: between filter_value: 24m # 24 months ago filter_value_end: 1d # yesterday description: "Table only contains last 24 months" ``` **Time formats:** - `24m` - 24 months - `1y` - 1 year - `30d` - 30 days - `1d` - 1 day ## List Partitions Partition by specific values. ```yaml partitions: - dimension: Region predicate: in_list filter_value: us-east, us-west, europe description: "Table only has 3 regions" ``` ## Multiple Partitions Define multiple partition constraints: ```yaml partitions: - dimension: Order Date predicate: between filter_value: 12m filter_value_end: 1d description: "Last 12 months" - dimension: Region predicate: in_list filter_value: us-east, us-west description: "US regions only" ``` ## How It Works When a spec arrives: 1. Planner checks partition constraints 2. If query matches partition, uses this table 3. If query doesn't match, uses alternative table (if available) 4. Routes query to appropriate table based on constraints ## Use Cases ### Rolling Window Table ```yaml name: Recent Orders partitions: - dimension: Order Date predicate: between filter_value: 12m filter_value_end: 1d description: "Last 12 months of data" ``` ### Regional Partition ```yaml name: US Sales partitions: - dimension: Region predicate: in_list filter_value: us-east, us-west, us-central description: "US regions only" ``` ## Best Practices 1. **Document partitions** - Add descriptions explaining constraints 2. **Use appropriate dimensions** - Name a dimension by name; it does not have to live on this table, but it must exist in the branch 3. **Set appropriate costs** - Partitioned tables often have lower cost 4. **Test query routing** - Verify planner routes queries correctly ## Next Steps - Learn about [cost optimization](#cost-optimization) - Read how [semantic routing](#semantic-routing) uses partitions ### Cardinality Understand relationship cardinality types. **danger: Many-to-Many Not Supported** 0sql does not support `many_to_many` relationships. Use junction/bridge tables with two separate relationships (`one_to_many` and `many_to_one`) instead. ## Overview Cardinality defines the relationship type between two tables. It tells the planner how rows in one table relate to rows in another. ## Cardinality Types ### many_to_one Most common type. Many rows in left table match one row in right table. **Example:** Many orders belong to one customer ```yaml orders_customer: left: Orders right: Customer sql: left.customer_id = right.id cardinality: many_to_one ``` **Use cases:** - Fact table to dimension table - Child to parent relationships - Foreign key relationships ### one_to_many One row in left table matches many rows in right table. **Example:** One customer has many orders ```yaml customer_orders: left: Customer right: Orders sql: left.id = right.customer_id cardinality: one_to_many ``` **Use cases:** - Parent to child relationships - One-to-many associations - Reverse of many_to_one ### one_to_one One row in left table matches one row in right table. **Example:** One user has one profile ```yaml user_profile: left: User right: User Profile sql: left.id = right.user_id cardinality: one_to_one ``` **Use cases:** - User profiles - Settings tables - Extension tables ## many_to_many Not Supported ⚠️ 0sql does not support `many_to_many` relationships. **Workaround:** Create a junction/bridge table and use two relationships: ```yaml # Instead of many_to_many: # Users <-> Roles # Create: # Users -> UserRoles -> Roles users_user_roles: left: Users right: User Roles sql: left.id = right.user_id cardinality: one_to_many user_roles_roles: left: User Roles right: Roles sql: left.role_id = right.id cardinality: many_to_one ``` ## Choosing Cardinality ### Fact to Dimension (many_to_one) ```yaml sales_customer: left: Sales # Fact table (many rows) right: Customer # Dimension table (one row per customer) sql: left.customer_id = right.id cardinality: many_to_one ``` ### Dimension to Dimension (one_to_one) ```yaml customer_address: left: Customer # One customer right: Address # One address per customer sql: left.address_id = right.id cardinality: one_to_one ``` ### Parent to Child (one_to_many) ```yaml customer_orders: left: Customer # One customer right: Orders # Many orders per customer sql: left.id = right.customer_id cardinality: one_to_many ``` ## Measure Expansion **warning: Understanding allow_measure_expansion** By default, measures can only be aggregated along `one_to_one` join paths. The `allow_measure_expansion` property enables measures to aggregate along `many_to_one` or `one_to_many` paths, but **use with caution** to avoid double-counting. ### The Double-Counting Problem When you join a fact table to a dimension through a `many_to_one` relationship and then to another table, measures might be counted multiple times: ``` Orders (fact) → Customer → Customer Address 100 rows 50 rows 50 rows ``` If you query `Total Revenue` grouped by `Address City`, each order amount could be counted once (correct) or multiple times (incorrect) depending on the join path. ### Using allow_measure_expansion Set `allow_measure_expansion: true` when you're **certain** the join path won't cause double-counting: ```yaml # Safe: Customer has exactly one address customer_address: left: Customer right: Customer Address sql: left.address_id = right.id cardinality: many_to_one allow_measure_expansion: true # Measures can flow through this join ``` ### When to Use | Scenario | allow_measure_expansion | Why | |----------|------------------------|-----| | Fact → Dimension (standard) | Not needed | Default behavior handles this | | Dimension → Dimension (1:1 logical) | `true` | Safe when truly 1:1 in practice | | Dimension → Lookup table | `true` | Safe for simple lookups | | Any path that could fan out | `false` (default) | Prevents double-counting | ### Example: Safe Expansion ```yaml datasource: warehouse # Orders to Customer: standard fact-to-dimension orders_customer: left: Orders right: Customer sql: left.customer_id = right.id cardinality: many_to_one # Customer to Address: each customer has one address # Safe to expand measures through this join customer_address: left: Customer right: Customer Address sql: left.address_id = right.ca_address_sk cardinality: many_to_one allow_measure_expansion: true ``` With this configuration, you can query `Total Revenue` by `Address City` and get correct results. ### Automatic Behavior - `one_to_one` relationships: Measures always expand (bidirectional) - `many_to_one` without flag: Only dimensions expand, measures stop - `many_to_one` with `allow_measure_expansion: true`: Measures expand through ## Internal Behavior Notes **info: How 0sql Handles Cardinality** - **one_to_many** is automatically converted to **many_to_one** internally (direction flipped) - **one_to_one** creates bidirectional joins automatically, so you can traverse either direction - Tables in a relationship must be in the same datasource ## Best Practices 1. **Use many_to_one** for fact-to-dimension relationships 2. **Match database relationships** - Cardinality should reflect actual data 3. **Avoid many_to_many** - Use junction tables instead 4. **Use allow_measure_expansion sparingly** - Only when you're certain it won't cause double-counting 5. **Document relationships** - Add comments explaining why expansion is safe ## Next Steps - [Join types](#join-types) - [Relationships guide](#modeling-with-zsql) ### Join Types Specify join types for relationships. ## Overview By default, relationships use INNER joins. You can specify different join types to handle optional relationships. ## Join Type Options ### inner (Default) Returns only rows that match in both tables. ```yaml orders_customer: left: Orders right: Customer join: inner # Optional, this is the default sql: left.customer_id = right.id cardinality: many_to_one ``` **Use cases:** - Required relationships - Most common case - Default if not specified ### left Returns all rows from left table, NULLs for non-matching right rows. ```yaml store_returns_reason: left: Store Returns right: Reason join: left sql: left.sr_reason_sk = right.r_reason_sk cardinality: many_to_one ``` **Use cases:** - Optional relationships - Nullable foreign keys - When right table may not exist ### right Returns all rows from right table, NULLs for non-matching left rows. ```yaml products_categories: left: Products right: Categories join: right sql: left.category_id = right.id cardinality: many_to_one ``` **Use cases:** - Less common - When you want all categories, even without products **note: Full Outer Join Not Supported** 0sql accepts only `inner`, `left`, and `right` as a relationship `join` value. `full` (FULL OUTER) is not valid and fails validation at `zsql check` or `zsql deploy`. If you need all rows from both tables regardless of match, consider: - Using two separate LEFT joins and combining results (e.g. in a view or downstream) - Restructuring your model to avoid the need for full outer joins ## Common Patterns ### Required Relationship (INNER) ```yaml sales_customer: left: Sales right: Customer sql: left.customer_id = right.id cardinality: many_to_one # join: inner (default, not needed) ``` ### Optional Relationship (LEFT) ```yaml orders_promotion: left: Orders right: Promotion join: left sql: left.promotion_id = right.id cardinality: many_to_one ``` **Why LEFT?** Not all orders have promotions, but we want all orders. ## Best Practices 1. **Use INNER by default** - Most relationships are required 2. **Use LEFT for optional** - When right table may not have matches 3. **Use RIGHT sparingly** - Less common, can be confusing 4. **Document join types** - Add comments explaining why ## Next Steps - Learn about [cardinality](#cardinality) - Read the [relationships guide](#modeling-with-zsql) - Explore [relationship schema reference](#cardinality) ### Tables Table models define semantic metadata for physical database tables. ## Overview Table models (`tbl.*.yml`) are the foundation of your semantic model. They map physical database tables to business-friendly semantic definitions. ## Table Structure ```yaml name: Store Sales physical_name: store_sales datasource: tpcds cost: 100 fields: - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` ## Required Fields ### name Logical name for the table in the semantic model. Must be unique within the datasource. ```yaml name: Store Sales ``` **Notes:** - Used in relationships and in the `tables` discovery endpoint - Never appears in a query spec: specs name fields, and the planner picks the table - Can differ from physical table name ### physical_name The actual table name in the database. Can include schema/catalog prefix. ```yaml physical_name: store_sales # Or with schema: physical_name: sales.orders ``` ### datasource Datasource key or name this table belongs to. ```yaml datasource: tpcds ``` ### cost Numeric value influencing table selection. Lower cost = preferred. ```yaml cost: 100 ``` **Guidelines:** - Dimension tables: `cost: 10` (lower) - Fact tables: `cost: 100` (higher) - Hot tier: Lower cost - Cold tier: Higher cost ## Optional Fields ### snapshot For snapshot tables (e.g., inventory), name the date dimension the table is snapped by. The dimension must be a date type, on this table or unambiguous in its universe. ```yaml snapshot: Date ``` Enables snapshot measures on the table (`snapshot: beginning` / `snapshot: ending` on a measure). The two keys share a name on purpose: the table's `snapshot` says *which date*, the measure's says *which end of the period*. ### tags Metadata tags for organization. Stored on the table; nothing in planning reads them. ```yaml tags: - marketing - operations - finance ``` ### partitions Data availability constraints that rank the table for routing. Predicates that rank: `between`, `greater_than`, `greater_than_or_equal_to`, `less_than`, `less_than_or_equal_to` (number or date dimensions) and `in_list` (any dimension). `partition:` is accepted as an alias. ```yaml partitions: - dimension: Order Date predicate: between filter_value: 24m # 24 months ago filter_value_end: 1d # yesterday ``` **Use cases:** - Rolling window tables - Partitioned tables - Region-specific data ### imports Import field definitions from other files. ```yaml imports: - ../common/common_dates.yml - ../common/geo_dimensions.yml ``` **Notes:** - Imports processed first - Local fields merged after - Supports relative paths - Validated for circular dependencies ## Fields Define dimensions and measures for the table. ```yaml fields: - type: dimension name: Customer Name data_type: string expression: sql: customer_name - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` [Learn about fields →](#dimensions) ## File Organization ### Naming Convention - File pattern: `tbl.{table_name}.yml` - Examples: - `tbl.orders.yml` - `tbl.store_sales.yml` - `tbl.dse.call_center_d.yml` (with schema) ### Directory Structure Organize tables by domain. `zsql new table NAME --datasource DS` writes `models//tbl..yml` (default domain `core`); any depth under `models/` is found. ``` models/ ├── sales/ │ ├── tbl.orders.yml │ ├── tbl.customers.yml │ └── rel.sales.yml ├── inventory/ │ ├── tbl.products.yml │ └── tbl.stock.yml └── common/ ├── tbl.date_dim.yml └── tbl.geography.yml ``` ## Complete Example ```yaml name: Store Sales physical_name: store_sales datasource: tpcds cost: 100 # Optional: Import common fields imports: - ../common/common_dates.yml # Optional: Define partitions partitions: - dimension: Order Date predicate: between filter_value: 24m filter_value_end: 1d fields: - type: dimension name: Store Ticket number description: Actual ticket number of an order data_type: integer expression: primary_key: true sql: ss_ticket_number - type: measure name: Store Quantity description: Quantity of items sold data_type: integer expression: sql: sum(ss_quantity) - type: measure name: Store Sales Price description: Actual sales price data_type: decimal format: currency:2 synonyms: - sales price - selling price - store revenue expression: sql: sum(ss_sales_price) ``` ## Best Practices 1. **Use descriptive names** - Business-friendly table names 2. **Set appropriate costs** - Lower for dimensions, higher for facts 3. **Organize by domain** - Group related tables 4. **Import common fields** - Reuse field definitions 5. **Define partitions** - Help planner route queries 6. **Add descriptions** - Document table purpose 7. **Set a cost** - A table without one is assumed to cost 0, with a warning from `zsql check` ## Next Steps - Learn about [dimensions](#dimensions) - Learn about [measures](#measures) - Understand [expressions](#sql-expressions) - [Imports](#imports) for reusing fields --- ## Advanced Features In-depth and optional semantic-model features. The pages above this section cover the core: tables, fields, expressions, relationships, imports, extended blending groups, and partitions. ## Features - **[Universe Formation](#universe-formation)** - How join trees are built from a root fact - **[Semantic Routing](#semantic-routing)** - How the planner picks a table and datasource - **[Cost Optimization](#cost-optimization)** - Steering routing with cost - **[Exclusions](#exclusions)** - Measures that ignore chosen dimensions - **[Inclusions](#inclusions)** - Multi-level calculations - **[Decorators](#projections-and-decorators)** - Request-time transforms on a projection (temporal, window, contribution). Nothing to model; they live in the query spec ## When to Use Use advanced features when you need: - **Point-in-time analysis** - [Snapshot measures](#snapshot-measures) - **Complex filtering** - Exclusions - **Multi-level calculations** - Inclusions - **Hot tiers and aggregate tables** - Semantic routing and cost optimization - **Time-based analysis** - [Temporal decorators](#projections-and-decorators) in the spec ## Next Steps - Explore specific advanced features - Read about [decorators](#projections-and-decorators) in the query spec - Learn how [semantic routing](#semantic-routing) picks a datasource ### Cost Optimization Optimize query performance with cost-based routing. ## Overview Cost optimization helps the planner choose the best table when multiple tables can answer the same question. Lower cost = preferred. ## Table Cost Set cost on each table: ```yaml name: Store Sales cost: 100 # Higher cost = used when needed ``` ```yaml name: Date Dimension cost: 10 # Lower cost = preferred ``` ## Cost Guidelines ### Dimension Tables Lower cost (preferred): ```yaml name: Customer cost: 10 # Dimension tables: 10 ``` ### Fact Tables Higher cost (used when needed): ```yaml name: Sales cost: 100 # Fact tables: 100 ``` ### Tier-Based Cost Match cost to datasource tier: ```yaml # Hot tier datasource warehouse: tier: hot # Tables in hot tier: cost 10-50 # Cold tier datasource archive: tier: cold # Tables in cold tier: cost 200-500 ``` ## Cost Ranges **Recommended ranges:** - Dimension tables: `10` - Fact tables: `100` - Hot tier: `10-50` - Warm tier: `50-100` - Cold tier: `100-500` ## How It Works When multiple tables can answer a query: 1. Planner identifies all possible tables 2. Checks partition constraints 3. Selects table with lowest cost 4. Routes query to selected table ## Examples ### Star Schema ```yaml # Dimension (low cost) name: Customer cost: 10 # Dimension (low cost) name: Product cost: 10 # Fact (high cost) name: Sales cost: 100 ``` ### Multi-Tier ```yaml # Hot tier table name: Recent Sales datasource: hot_warehouse cost: 10 # Cold tier table name: Historical Sales datasource: cold_archive cost: 200 ``` ## Best Practices 1. **Set appropriate costs** - Lower for dimensions, higher for facts 2. **Match tier to cost** - Hot tier = lower cost 3. **Use partitions** - Help planner route queries 4. **Test query routing** - Verify planner selects correct table ## Next Steps - Learn about [partitions](#partitions) - Read about [semantic routing](#semantic-routing) - Read about [explain](#explain) to see which table a spec resolved to ### Exclusions Exclude dimensions from measure calculations. ## Overview Exclusions let you control **how a measure groups and responds to filters** on specific dimensions, tables, or whole universes. There are two independent knobs: - **Exclusion type (`exclusion_type`)** controls **grouping**: which dimensions are allowed to affect the grain of the measure. - **Filter behavior (`filter`)** controls **filters**: whether filters on those entities are applied, ignored, or treated as the only filters that matter for this measure. Grouping and filtering are independent of each other. Keeping that in mind is the key to modeling these measures correctly. ## Exclusion Types (grouping behavior) ### `exclude` Exclude specific entities from grouping while leaving all other dimensions available to group the measure. ```yaml - type: measure name: Revenue (Excluding Returns) data_type: decimal exclusion_type: exclude exclusions: - type: dimension filter: apply entities: - Return Status expression: sql: sum(amount) ``` In this example, `Return Status` is left out of the grouping for this measure even if the user adds it to the query, while other dimensions such as `Date` or `Customer` still group it. ### `exclude_all_except` Allow **only** the listed entities to affect grouping for this measure. Every other dimension is excluded from grouping. ```yaml - type: measure name: Revenue (Only by Product) data_type: decimal exclusion_type: exclude_all_except exclusions: - type: dimension filter: apply entities: - Product Category expression: sql: sum(amount) ``` Here the measure always behaves as if it were aggregated by `Product Category` alone, even when more dimensions are present in the query. ### `exclude_all` Exclude every dimension from grouping for this measure. The measure is computed as one global value, regardless of what the query groups by. ```yaml - type: measure name: Total Revenue (No Grouping) data_type: decimal exclusion_type: exclude_all expression: sql: sum(amount) ``` This is useful for overall-total or grand-total style measures. ## Exclusion Structure ```yaml exclusion_type: exclude # exclude, exclude_all_except, exclude_all exclusions: - type: dimension # dimension, table, universe filter: apply # ignore, only, apply entities: - Dimension Name ``` ### Entity types (`type`) - **`dimension`**: Match dimensions by name. Only the listed dimensions are affected. - **`table`**: Apply the rule to **all dimensions that come from the given table**. Useful for switching off an entire hierarchy, such as every product attribute at once. - **`universe`**: Apply the rule to **all dimensions reachable from a given universe root**. Powerful, so use it deliberately: it can affect many dimensions at once. Use `dimension` by default, and only escalate to `table` or `universe` when you have a clear, intentional modeling need. ## Filter Options (filter behavior) Filter behavior controls how **filters on the target entities** are treated for this measure. It does **not** change which dimensions can group the measure; that is controlled by `exclusion_type`. ### `ignore` Ignore filters on the target entities for this measure. Grouping still follows `exclusion_type`. ```yaml filter: ignore ``` **Important**: This controls **filters**, not visibility. The target dimension can still appear in the result as a row (depending on `exclusion_type`), but any filter applied to it is **ignored** when this measure is computed. Example: a "% of total revenue" measure can ignore category filters, so that each row shows revenue as a share of the **unfiltered** total even when the user has filtered to one category. ### `only` Apply **only** the filters on the target entities to this measure. Filters on every other dimension are ignored for this measure, even though those dimensions can still group it. ```yaml filter: only ``` Example: a measure that should always respect `Product Category` filters but ignore any ad hoc filters on `Customer Segment` or `Region`. ### `apply` Apply filters on the target entities normally. This is the default behavior. ```yaml filter: apply ``` **Important**: The filters are applied, but the dimension may still be **excluded from grouping**, depending on `exclusion_type`. The dimension can then appear in the result as rows while the measure is computed at a coarser grain, so the filtered value repeats across those rows. ## Examples ### Example 1: % of Total Revenue by Category **Goal:** Show each product category's revenue and its share of **overall** revenue, even when filters are applied. ```yaml - type: measure name: Revenue data_type: decimal expression: sql: sum(amount) - type: measure name: Revenue (All Categories) data_type: decimal exclusion_type: exclude exclusions: - type: dimension filter: ignore # ignore category filters entities: - Product Category expression: sql: sum(amount) - type: measure name: Revenue % of Total data_type: decimal format: percent:2 expression: sql: "[Revenue] / nullif([Revenue (All Categories)], 0)" ``` - Grouping: `Product Category` groups `Revenue`, but it is excluded from `Revenue (All Categories)`, so that measure is the total across all categories, repeated on every category row. - Filters: filters on `Product Category` are **ignored** for `Revenue (All Categories)` because of `filter: ignore`, so the denominator is always the unfiltered total. ### Example 2: Measure that Ignores a Whole Table **Goal:** Compute revenue that is unaffected by any product-level filter, whichever product dimensions are used. ```yaml - type: measure name: Revenue (Ignoring Product Hierarchy) data_type: decimal exclusion_type: exclude exclusions: - type: table filter: ignore entities: - Product # table display name expression: sql: sum(amount) ``` - Grouping: product dimensions still appear in the result set, but they do not change this measure's grain. - Filters: any filter on a dimension from the `Product` table is ignored for this measure. ### Example 3: Dimension Present vs Absent in the Query **Goal:** Show how an exclusion behaves depending on whether the excluded dimension is projected. **Setup:** ```yaml - type: measure name: Revenue (Excluding Category) data_type: decimal exclusion_type: exclude exclusions: - type: dimension filter: apply entities: - Product Category expression: sql: sum(amount) ``` **Case A: Category IS projected** - **Query**: Revenue (Excluding Category) grouped by Date, Product Category - **Behavior**: `Product Category` appears as rows in the result, but the measure is computed **without the category split**. It is aggregated at the Date level, and that Date total is repeated on each category row. - **Result**: every category row for a given date shows the **same value**, the Date-level aggregate. | Date | Product Category | Revenue (Excluding Category) | |------|-----------------|------------------------------| | 2024-01-01 | Electronics | $10,000 | | 2024-01-01 | Footwear | $10,000 | | 2024-01-01 | Apparel | $10,000 | **Case B: Category IS NOT projected** - **Query**: Revenue (Excluding Category) grouped by Date only - **Behavior**: the measure is computed at the Date level, since nothing else groups it. - **Result**: one value per Date, the same total that Case A repeats on every category row. | Date | Revenue (Excluding Category) | |------|------------------------------| | 2024-01-01 | $10,000 | **Key insight**: an exclusion changes the **aggregation level**, not whether the dimension is displayed. When the excluded dimension is projected you still get it as rows, but the measure **ignores it for grouping**. This is how "% of total" measures work: every row carries the same denominator, the grand total, so `[Category Revenue] / [Revenue (Excluding Category)]` gives each category's share. ## Best Practices 1. **Use descriptive names**: "Revenue (Excluding Returns)", not just "Revenue". 2. **Document exclusions**: add a description saying what is excluded and why. 3. **Start with the `dimension` type**: use `table` or `universe` only when you need that breadth. 4. **Test carefully**: compare results with and without the exclusion, and check the edge cases (no filters, several filters, dimensions added and removed). 5. **Use sparingly**: exclusions are powerful but easy to misread; prefer a simpler measure when one will do. ## Next Steps - Learn about [inclusions](#inclusions) - Explore [snapshot measures](#snapshot-measures) - Read about the [`contribute` decorator](#projections-and-decorators), the request-time way to get a percent of total ### Inclusions Force multi-level calculations with dimension inclusions. ## Overview Inclusions solve a critical problem: **computing correct averages and medians when counts per group are uneven**. They force calculations at a finer grain (by including extra dimensions) before rolling up to the final result. This is the opposite direction to [exclusions](#exclusions), which remove dimensions from calculations. Inclusion measures enable multi-level aggregations like medians, percentiles, and other complex calculations that require intermediate grouping. ## What Problem Does This Solve? ### Example: Daily Median Account View Hours Imagine you have a table tracking video viewing activity at the grain of `(day, account, title)`. Each row represents how many hours a specific account watched a specific title on a specific day. You want to compute: **Measure**: Daily median account view hours **Wrong approach:** ```yaml - type: measure name: Median View Hours (Wrong) data_type: decimal expression: sql: median(hours) # Computed at (day, account, title) grain ``` This produces a meaningless "account-title median" because you're taking the median across **every row** (every title view), not across **accounts**. **What you actually want:** 1. First, **sum** hours per account per day (collapse the title dimension) 2. Then, compute the **median** of those account totals for each day **Correct approach with inclusion:** ```yaml - type: measure name: Daily Median Account View Hours data_type: decimal inclusions: filter: apply aggregation: percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp) dimensions: - Account ID expression: sql: sum(hours) ``` **How this works:** - **Inner aggregation**: `SUM(hours)` per `(day, account)` → total hours each account watched each day - **Outer aggregation**: `MEDIAN(...)` per `day` → median across all accounts for that day This ensures you're computing the median of **account totals**, not individual title-view rows. ### When You Need Inclusions Use inclusion measures when: - Computing **averages** or **medians** where the underlying data has uneven counts per group - You need to aggregate at an **intermediate grain** (e.g., per account) before computing a final statistic (e.g., median across accounts) - Standard aggregation would produce incorrect results due to grain mismatch ## How It Works Inclusion measures produce a **two-step SQL shape**: ### Step 1: Inner Aggregation (with included dimensions) The engine first computes the base expression (`sql: ...`) grouped by: - **Target dimensions** (the dimensions in the user's query, e.g., `Date`) - **Plus included dimensions** (the dimensions listed in `inclusions.dimensions`, e.g., `Account ID`) ```sql -- Inner query SELECT day, account_id, sum(hours) AS account_hours -- base expression FROM viewing_activity GROUP BY day, account_id ``` ### Step 2: Outer Aggregation (without included dimensions) The engine then applies the `aggregation` function to the results of step 1, grouped by: - **Target dimensions only** (included dimensions are dropped) ```sql -- Outer query SELECT day, percentile_cont(0.5) WITHIN GROUP (ORDER BY account_hours) AS median_hours FROM ( -- Inner query (from step 1) SELECT day, account_id, sum(hours) AS account_hours FROM viewing_activity GROUP BY day, account_id ) subq GROUP BY day ``` The `@exp` placeholder in the `aggregation` field is replaced with the result of the inner aggregation (e.g., `account_hours` in the example above). ### SQL Summary | Step | GROUP BY | Aggregation | |------|----------|-------------| | Inner | Target dims + Included dims | Base expression (`sql`) | | Outer | Target dims only | Outer aggregation (`aggregation`) | ## Industry Context If you're familiar with other BI or semantic layer tools, 0sql's inclusion measures are similar to: - **Tableau**: INCLUDE [Level of Detail (LOD) calculations](https://help.tableau.com/current/pro/desktop/en-us/calculations_calculatedfields_lod.htm). The INCLUDE keyword adds dimensions to the calculation grain. - **MicroStrategy**: Has an equivalent feature for level-based metrics (specific feature name varies by version). - **Looker**: Does **not** support this pattern natively. You would need to pre-build derived tables with the intermediate aggregation. The key insight across all these tools: when your raw data is at a fine grain (e.g., per title view) but you need to aggregate at an intermediate grain (e.g., per account) before computing a final statistic, you need a way to specify that intermediate grouping explicitly. ## Configuration ```yaml inclusions: filter: apply # ignore, only, apply aggregation: percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp) # Outer aggregation dimensions: - Account ID # Dimensions to include in inner aggregation ``` ### Configuration Fields **`dimensions`**: List of dimension names to temporarily include in the **inner aggregation** (step 1). These dimensions are added to the GROUP BY in the inner query, then dropped in the outer query. **`aggregation`**: SQL expression for the **outer aggregation** (step 2). Use `@exp` as a placeholder for the result of the inner aggregation. Common patterns: - `percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp)`: median - `percentile_cont(0.9) WITHIN GROUP (ORDER BY @exp)`: 90th percentile - `avg(@exp)`: average of the inner results - `stddev(@exp)`: standard deviation - `max(@exp)`, `min(@exp)`: max or min of the inner results **`filter`**: Controls how filters on the **inclusion dimensions** are treated in the inner aggregation: - `apply` (default): filters on inclusion dimensions are applied in the inner query - `ignore`: filters on inclusion dimensions are ignored in the inner query - `only`: only filters on inclusion dimensions are applied; other filters are ignored ## Use Cases ### Median Calculation ```yaml - type: measure name: Median Order Value data_type: decimal inclusions: filter: apply aggregation: percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp) dimensions: - Product Category expression: sql: sum(amount) ``` **Generated SQL:** 1. **Inner**: `sum(amount)` grouped by (query dims + Product Category) 2. **Outer**: `percentile_cont(0.5)` grouped by query dims only → median across Product Category values ### Percentile Calculation ```yaml - type: measure name: 90th Percentile Order Value data_type: decimal inclusions: filter: apply aggregation: percentile_cont(0.9) WITHIN GROUP (ORDER BY @exp) dimensions: - Customer Segment expression: sql: sum(amount) ``` ### Custom Aggregation ```yaml - type: measure name: Custom Measure data_type: decimal inclusions: filter: apply aggregation: stddev(@exp) # Standard deviation dimensions: - Region expression: sql: sum(amount) ``` ## Complete Example ```yaml - type: measure name: Median Revenue by Product description: Median revenue calculated at product level data_type: decimal format: currency:2 inclusions: filter: apply aggregation: percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp) dimensions: - Product Category expression: sql: sum(amount) ``` ## Limitations **Multiple inclusion rules**: `inclusions` is a single mapping, not a list. The engine handles **one inclusion rule per measure**. If you need multiple intermediate groupings, consider: - Defining separate measures for each grouping level, or - Using a single inclusion with the most critical intermediate dimension ## Best Practices 1. **Use for complex calculations** - Medians, percentiles, standard deviation, averages with uneven counts 2. **Document aggregation** - Explain what the inclusion does and which dimension(s) provide the intermediate grain 3. **Test carefully** - Verify multi-level calculations work correctly; compare against hand-written SQL 4. **Use descriptive names** - "Daily Median Account View Hours" not "Median Hours" 5. **One inclusion per measure** - Stick to a single inclusion rule for predictable behavior ## Next Steps - Learn about [exclusions](#exclusions) - Explore [snapshot measures](#snapshot-measures) - Read about [window decorators](#projections-and-decorators) in the query spec ### Semantic Routing What you configure so the query planner picks the right table and datasource. ## Overview When a query requests measures and dimensions, the planner chooses which **table** (and thus **datasource**) to use. You don't hard-code table names, you set **tier**, **cost**, and **partitions** in YAML, and the planner uses those to route queries. This page is about what each setting means and where to set it. **Rule you must respect:** a spec is planned against exactly one datasource. 0sql never blends across datasources, so everything a spec asks for must be servable from a single one. When the same measures and dimensions are defined in more than one datasource, routing decides which one the SQL is written for; the response names it in `datasource` and `adapter`. ## What You Configure ### Datasource Tier **Where:** In your datasource config (e.g. `datasources.yml`). **What it means:** When more than one datasource can answer the query, the planner prefers the one with the **lower tier number**: hot over warm over cold. | Tier | Typical use | |------|-------------| | `hot` | Primary warehouse, fast queries | | `warm` | Read replicas, moderate speed | | `cold` | Archives, data lakes, slow | ```yaml # datasources.yml primary: adapter: snowflake tier: hot replica: adapter: postgres tier: warm archive: adapter: athena tier: cold ``` **Example:** If the same data exists in hot (last 3 months) and cold (full history), a query for "last 30 days" will use the hot datasource; a query for "year 2020" will use cold. ### Table Cost **Where:** In each table's YAML (`cost` on the table). **What it means:** When several tables can answer the same query, the planner prefers **lower cost**. Use cost to prefer pre-aggregated or summary tables over detail tables. ```yaml # Fact / detail table: higher cost name: Store Sales cost: 100 # Pre-aggregated table: lower cost, preferred when it can answer the query name: Store Sales Daily cost: 20 ``` When two tables define the same measure (e.g. "Store Sales Price"), the [one-name-one-concept rule](#one-name-one-concept-rule) makes them one measure with two sources. The planner first keeps the tables whose dimensions cover the query, then prefers the lower-cost one, so `cost` is how you break ties. ### Partition Constraints **Where:** In the table's YAML under `partitions` (a list; the singular `partition` key is also accepted). **What it means:** This table holds data for a specific range (e.g. "Sale Date ≥ 2024-01-01"). The planner prefers tables whose partitions match the spec's filters and avoids tables whose partitions exclude them. The predicates that rank a table are `between`, `greater_than`, `greater_than_or_equal_to`, `less_than`, `less_than_or_equal_to` (number or date dimensions) and `in_list` (any dimension). See [Partitions](#partitions) for full syntax. ```yaml name: Recent Sales partitions: - dimension: Sale Date predicate: greater_than_or_equal_to filter_value: 2024-01-01 ``` ## Dimension and Measure Coverage The chosen table's **universe** must contain paths to **all** requested dimensions and the requested measures. You don't configure this per-query, you design your semantic model (joins, cardinality) so that the right universe exists. If you see errors like "dimension not in universe," fix joins or cardinality so that one universe can reach all requested fields. See [Universe Formation](#universe-formation). ## Best Practices 1. **Use meaningful costs**: Pre-aggregated tables should have lower costs 2. **Configure tiers correctly**: Hot for fast, cold for archives 3. **Add partitions**: Help planner route time-filtered queries 4. **Test routing**: `zsql explain --expr "..."` prints the datasource and the table route the planner chose for a spec 5. **Document tier boundaries**: Know what data each tier contains ## Next Steps - [Universe Formation](#universe-formation): How paths are generated - [Cost Optimization](#cost-optimization): Fine-tuning costs - [Partitions](#partitions): Time-based constraints ### Universe Formation Understanding how 0sql builds query paths through your semantic model. ## Overview When you run `zsql deploy`, the **Formation Engine** on the 0sql service generates all possible paths from every table to every reachable field. These paths form **universes**: complete, cardinality-safe data access routes that the planner uses to turn a spec into SQL. Understanding universe formation helps you: - **Design efficient semantic models** – keep join graphs sane and avoid double-counting - **Understand query routing** – why a particular table/universe was chosen for a query - **Reason about dimensionality** – which dimensions are allowed to group a given measure - **Validate measure eligibility** – why some combinations of measures/dimensions are invalid If you are familiar with other tools, a 0sql universe plays a role similar to a **BusinessObjects Universe** or a **Looker Explore**, but it is derived automatically from your YAML model and enforced with strict cardinality and routing rules. ## What is a Universe? A **universe** is a set of paths from a single root table to all fields reachable via joins. Each table in your model has its own universe. ```mermaid flowchart TD subgraph StoreSalesUniverse["Store Sales Universe"] SS["Store Sales (Root Table)"] SS --> D["Date Dimension"] SS --> C["Customer Dimension"] SS --> S["Store Dimension"] C --> CA["Customer Address"] CA --> ST["State Lookup"] end ``` The Store Sales universe contains paths to: - All fields directly on Store Sales - All fields on Date, Customer, Store (via direct joins) - All fields on Customer Address (via Customer) - All fields on State Lookup (via Customer → Customer Address) ## How Universes Are Built On deploy, the Formation Engine runs on the service and discovers all paths from each table to every reachable field by following your join definitions. It applies cardinality and measure-expansion rules so that only valid, cardinality-safe paths are available to the planner. You don't configure paths directly: they are derived from your semantic model. `zsql deploy` reports the path count (`N paths`) and any formation warnings, such as two routes of equal cost. ### What You Need to Know: Measure Expansion Not all paths allow measure aggregation. When you design joins, these rules determine whether measures can flow through a path (and thus whether a measure can be combined with dimensions on the other side of that join): | Join Type | Measures Expand? | Why | |-----------|-----------------|-----| | `one_to_one` | Always | No fan-out risk | | `many_to_one` with `allow_measure_expansion: true` | Yes | Explicitly allowed | | `many_to_one` without flag | No | Prevents double-counting | | `one_to_many` | No | Would multiply measures | ```yaml # This path allows measures to flow through: customer_address: cardinality: many_to_one allow_measure_expansion: true # This path blocks measures (dimensions only): customer_orders: cardinality: one_to_many # Measures from Customer cannot aggregate across Orders ``` Getting cardinality and measure expansion right is what matters when you design your semantic model; the rest is handled by the engine. ## Universe Selection When a spec arrives with specific measures and dimensions, the planner selects which universe(s) can answer it: a universe must contain paths to **all** requested dimensions and **any** requested measures. When several universes qualify, the planner ranks them (e.g. by tier, cost, partition fit) and picks the best fit. Details are in [Semantic Routing](#semantic-routing). **Example:** For a query like *Total Revenue by Customer State*, the Store Sales universe is used because it has the measure (Total Revenue) and can reach the dimension (Customer State via Customer → Customer Address). The Customer universe is not used because it doesn’t contain that measure. If you see errors like "dimension not in universe," it means no single universe has a path to all requested fields, often a join or cardinality issue in your model. ## Path Cost Calculation Path cost influences table selection when multiple tables can answer the same query: A join's cost is the sum of the two tables' `cost` values, and a path's cost is the sum of the joins it takes. Cardinality does not change the cost; it decides whether measures may travel the path at all. ```yaml # Table costs (set in tbl.*.yml) Store Sales: cost: 100 # Fact table - higher Customer: cost: 10 # Dimension - lower # The Store Sales -> Customer join therefore costs 110 ``` Lower total cost = preferred path. ## Blended Dimensions Dimensions with the same `extended_blend_group` can be used interchangeably: ```yaml # In tbl.store_sales.yml - name: Sale Date extended_blend_group: date_dimension # In tbl.catalog_sales.yml - name: Order Date extended_blend_group: date_dimension ``` The engine creates **blend paths** that allow queries to: 1. Use "Sale Date" when querying Store Sales measures 2. Use "Order Date" when querying Catalog Sales measures 3. Automatically blend when both measure types are requested ## Best Practices 1. **Keep join graphs simple**: Star schema is faster than complex snowflake 2. **Use appropriate cardinality**: Match actual data relationships 3. **Be careful with measure expansion**: Only enable when truly safe 4. **Use cost hints**: Lower cost on dimension tables, higher on facts 5. **Leverage tiers**: hot datasources are preferred when several can answer the spec ## Next Steps - [Semantic Routing](#semantic-routing): how a spec selects a table and datasource - [Cost Optimization](#cost-optimization): tuning table and join costs - [Explain](#explain): see the universe and paths a spec resolved to --- ## Security Security policies restrict which rows and values each caller can see. They live in `security.yml`, deploy with the model, and are applied by the planner, so the SQL that comes back already carries the WHERE clause or the CASE mask. 0sql has no end users of its own: your application decides who the caller is and sends that as the `context` of every request. Enforcement happens where the data is. 0sql never connects to your warehouse and never sees a row, so a policy is not a filter applied to results in transit: it is a predicate compiled into the statement before you run it. The rows a caller may not see are never read in the first place. ## How a policy works A policy has three parts, evaluated for every query node: 1. **Triggers** decide whether it fires: when the query projects a field carrying one of the policy's tags, or one of its named fields. A query that projects no triggered field is unaffected. 2. **Context dimension** is the boundary: the dimension whose values partition the data, such as `Call Center ID` or `Employee Email`. It must be reachable from every table that holds a triggered field; if it is not, the query is refused (422 `Planner::SecurityPolicyError`) rather than returning unprotected rows. 3. **Permission resolution** works out which values of that dimension the caller may see, from the groups or the user in the context. ### Modes | Mode | SQL that comes back | |---|---| | `filter_data` | `WHERE LOWER() IN ('a', 'b')` is added. Restricted rows are gone and aggregates count only what the caller may see. | | `mask_data` | Every row stays. Each triggered field becomes `CASE WHEN LOWER() IN ('a', 'b') THEN ELSE '' END`; numeric fields mask to `NULL`. | ### The context Your application sends the caller's security context in the request body `context`. The planner reads it and stores nothing. ```json { "email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": ["region:EMEA"], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}] } ``` Unknown keys are rejected. Tags are `key:value` strings, on the user (`tags`) or on each group. Where they come from (your identity provider, your own database, a JWT) is up to you; 0sql only sees what you send. A branch with at least one policy refuses to plan without a context (HTTP 400): ```json {"error": {"class": "ContextRequired", "message": "this branch has security policies; a context is required to plan"}} ``` On a branch without policies the context is optional. ## Tagging fields Add `tags` to any dimension or measure in a table file: ```yaml fields: - type: dimension name: Customer First Name tags: [pii] data_type: string expression: sql: c_first_name ``` Use one tag per sensitivity class (`pii`, `hipaa`, `finance`) so protecting a new field means tagging it, not editing the policy. ## security.yml reference ```yaml policies: - name: PII Call Center Masking mode: mask_data # mask_data | filter_data mask_value: "#######" # mask_data only, default "#######" triggers: field_tags: [pii] # any field tagged "pii" field_names: # optional: specific fields by name - Customer First Name context_dimension: Call Center ID permission_resolution: source: groups # groups | user value_from: tag # groups: name | tag / user: email | tag tag_key: call_center_id # required when value_from is "tag" unresolved: deny # deny | allow bypass: system_admin: true project_admin: false ``` The file is optional. When present it must have a top-level `policies` list, and it is the source of truth for the branch on every `zsql deploy`: `policies: []` removes every policy. Policies are branch-scoped; one deployed to `staging` does not affect `main`. ### name Required. The policy's identity within the branch. ### mode Required. `mask_data` or `filter_data`. ### mask_value `mask_data` only. The string shown in place of restricted string values. Defaults to `#######`. Numeric fields are masked to `NULL` regardless. ### triggers - **field_tags**: fires when a query projects any field carrying one of these tags. - **field_names**: explicit dimension or measure names, resolved at deploy time. A name that does not exist fails the deploy: `Policy '': trigger field '' not found in this branch.` A policy with neither never fires. ### context_dimension Required. A dimension name, resolved at deploy time (`Policy '': context_dimension '' not found in this branch.`). It may not be a composite bracket-reference expression. Every table holding a triggered field needs a join path to it, so pick a key that every fact carries. If a later deploy drops the dimension, the policy stays but is inert and the deploy warns `policy context dimension not found; policy is inert`. ### permission_resolution | source | value_from | Allowed values come from | |---|---|---| | `groups` | `tag` | Each group tag `tag_key:VALUE`. Groups tagged `call_center_id:TMNT` and `call_center_id:AVCL` allow both. | | `groups` | `name` | The group names. Name groups after the context values. | | `user` | `email` | The context `email`. Use with a dimension that holds emails. | | `user` | `tag` | Each user tag `tag_key:VALUE`. | `source` defaults to `groups`. `tag_key` is required when `value_from` is `tag`. Any other combination resolves nothing. **unresolved** decides what happens when a fired policy resolves no values. The default, `deny`, applies the policy with an empty set: `filter_data` returns no rows (`WHERE 1 = 0`) and `mask_data` masks every triggered value, so a caller your application forgot to tag sees nothing rather than everything. `allow` skips the policy for that caller, which suits a policy meant to restrict only some callers. ### bypass - **system_admin** (default `true`): a context with `"system_admin": true` skips the policy. - **project_admin** (default `false`): a context with `"project_admin": true` skips the policy. ## Examples ### Filter rows by group tag ```yaml policies: - name: Call center rows mode: filter_data triggers: field_tags: [pii] context_dimension: Call Center ID permission_resolution: source: groups value_from: tag tag_key: call_center_id unresolved: allow bypass: system_admin: true project_admin: false ``` ```json {"email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]}]} ``` Allowed values are `TMNT` and `AVCL`. A query projecting a `pii` field gains `WHERE LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl')`. A second policy keyed on `cat` adds its own `AND LOWER(...) IN ('men')`. A context with `"system_admin": true` gets the plain SELECT. ### Mask values by group name ```yaml policies: - name: Mask first name outside my call centers mode: mask_data mask_value: "#######" triggers: field_names: [Customer First Name] context_dimension: Call Center ID permission_resolution: source: groups value_from: name ``` ```json {"email": "tank@matrix.com", "groups": [{"name": "TMNT", "tags": []}, {"name": "AVCL", "tags": []}]} ``` Allowed values are the group names. Every row stays and `Customer First Name` is projected as `CASE WHEN LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl') THEN T0."c_first_name" ELSE '#######' END`. With `"groups": []` and the default `unresolved: deny` every value is masked. ### Filter rows to the caller's own record ```yaml policies: - name: Employee self service mode: filter_data triggers: field_tags: [employee_pii] context_dimension: Employee Email permission_resolution: source: user value_from: email bypass: system_admin: true project_admin: false ``` ```json {"email": "tank@matrix.com"} ``` The SQL gains `WHERE LOWER(employee_email) IN ('tank@matrix.com')`. The user-tag variant, `source: user`, `value_from: tag`, `tag_key: region`, is unlocked by `{"tags": ["region:EMEA"]}`. ## Testing a policy Put a context in a file and plan with it. ```sh cat > user.json <<'EOF' {"email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}]} EOF zsql sql --expr "customer first name, store net paid" --context user.json zsql explain --expr "customer first name, store net paid" --context user.json ``` `zsql sql` prints the SQL with the filter or mask in place. `zsql explain` adds the node tree; a node the policy touched lists it under `security_filters` or `security_masks`. Repeat with `"system_admin": true` to confirm the bypass, with `"groups": []` to see `unresolved` at work, and with no `--context` to get `ContextRequired`. Tests in `tests/*.yml` plan without a context, so they check the model, not the policies. ## Best practices 1. **Tag once, protect everywhere.** Prefer `field_tags` over `field_names`. 2. **Pick a context dimension every fact can reach.** An unreachable one fails the query, not the deploy. 3. **Keep `unresolved: deny`** and the system admin bypass on. 4. **Build the context in one place** in your application, so every call to `/sql` uses the same mapping from your identity provider to `tags` and `groups`. ## Next steps - [The security context in a request](#security-context) - [Explain output](#explain) - [Accounts, keys and access](#accounts-keys-and-access) --- ## Accounts and keys An account holds users, projects and keys. You deploy and administer with a personal key; your application queries with a read-only query key that is granted specific projects and branches. Nothing you send to plan a query is stored. ## Signing up Sign up at [https://app.0sql.io](https://app.0sql.io) with an account name, your name, email and a password of at least 8 characters. That creates the account, makes you its first user with the `admin` role, and shows your first personal key once: ``` Run `zsql auth --api-key ` inside a project to store it. You can make more keys under My keys. ``` Copy it. The console lists keys by their first 12 characters afterwards and never shows the secret again. Store it in the project with `zsql auth --api-key zsk_...`, which writes `.zsql` (mode 0600, git-ignored), or export `ZSQL_API_KEY`, which wins over any file. See [Authentication](#authentication-and-configuration). ## Roles | Role | Can | |---|---| | `developer` | Create personal and query keys, read every project with `visibility: account`, deploy where they hold Write, and create new projects (the first deploy of a new uid makes them its owner). | | `admin` | Everything a developer can, plus add and remove members, issue invites, revoke or rotate any query key, and read the audit log. | ## Two kinds of key | | Personal key | Query key | |---|---|---| | Prefix | `zsk_` | `zqk_` | | Belongs to | a user | the account | | Can do | whatever its user may do | read only: plan SQL, explain, explore, list fields and tables | | Scope | every project the user can reach | only the projects and branches it is granted | | Where it lives | `.zsql` on a developer machine, a CI secret | your application's configuration | | Rotation | revoke and create a new one | `rotate` issues a new secret and keeps the grants | Both go in the same header on every request: ```http Authorization: Bearer zsk_... ``` A key answers for one account. Secrets are stored hashed (SHA-256); the service cannot show them again. ### Personal keys Named, several per user, listed under **My keys** in the console with name, 12-character prefix, created and last-used times. `zsql` and CI use them to `deploy`, `check`, `test` and query. ```sh curl -s https://app.0sql.io/account/keys \ -H "Authorization: Bearer zsk_..." \ -H "Content-Type: application/json" \ -d '{"name": "ci"}' ``` Response: `{"key": {...}, "secret": "zsk_..."}`. Revoke with `DELETE /account/keys/{id}`. ### Query keys Always read only, even for an admin. A query key starts with no access; a project owner grants it a project, and the grant names a branch: | `branch` in the grant | Meaning | |---|---| | omitted (`null`) | the project's production branch | | `"*"` | every branch | | `"staging"` | that one branch | Create one: ```sh curl -s https://app.0sql.io/account/query-keys \ -H "Authorization: Bearer zsk_..." \ -H "Content-Type: application/json" \ -d '{"name": "checkout-service"}' ``` ```json {"key": {"id": "", "name": "checkout-service", "prefix": "zqk_abcdefgh", "created_by": "", "created_at": "...", "last_used_at": null}, "secret": "zqk_..."} ``` Grant it the `tpcds` project on every branch (you must own `tpcds`): ```sh curl -s https://app.0sql.io/account/query-keys//grants \ -H "Authorization: Bearer zsk_..." \ -H "Content-Type: application/json" \ -d '{"project": "tpcds", "branch": "*"}' ``` ```json {"key": {...}, "grant": {"project_id": "", "project_uid": "tpcds", "branch": "*"}} ``` Ship the secret with your application and plan with it: ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"expr": "item category, store net paid, year = 2002", "context": {"email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}]}}' ``` The `context` is the security context of the person your application is answering for. It is required on a branch that has policies and ignored on one that does not. See [Security](#row-level-security). `GET /me` tells a key what it is: ```json {"kind": "query_key", "key": {"id": "", "name": "checkout-service", "prefix": "zqk_abcdefgh", "created_by": "", "created_at": "...", "last_used_at": "..."}, "account": {"uid": "acme", "name": "Acme"}, "grants": [{"project_id": "", "project_uid": "tpcds", "branch": "*"}]} ``` A personal key gets `{"kind": "user", "user", "account", "role", "key_id"}` instead. Revoke a grant with `DELETE /account/query-keys/{id}/grants/{project}`. Rotate with `POST /account/query-keys/{id}/rotate`, which returns a new `secret` and keeps every grant; the old secret stops working. Revoke the key with `DELETE /account/query-keys/{id}`. Rotation and revocation are open to the key's creator and to admins. A query key can call only `GET /me`, `GET /projects` and the read routes of its granted branches: `sql`, `explain`, `explore`, `fields`, `tables` and the branch summary. Anything else is a 403. ### Which key for what | Task | Key | |---|---| | `zsql deploy`, `zsql check`, `zsql test` | personal | | `zsql sql`, `zsql explain`, `zsql fields` on a developer machine | personal | | CI pipeline that validates and deploys | personal (a dedicated user's key) | | An application calling `/sql` | query | | Granting access, changing project settings | personal, held by an owner | | Adding members, reading the audit log | personal, held by an admin | ## Projects A project is created by its first deploy: `zsql deploy` from a directory whose `project.yml` names a `uid` the account has not seen creates the project with that `uid`, `name` and `production_branch`, and makes the deploying user its owner. The uid is a slug: lower-case letters, digits, `-` and `_`. Each project has: | Setting | Values | Effect | |---|---|---| | `visibility` | `account` (default) \| `restricted` | `account`: every developer in the account has Read. `restricted`: only people granted a level. | | `protected_production` | `true` \| `false` | When true, Write on the production branch needs Owner. Other branches are unaffected. | | `production_branch` | branch name, default `main` | The branch a query key grant with no `branch` reads. | Owners change these with `PATCH /projects/{uid}` or the project's **Settings** tab, and delete a project with `DELETE /projects/{uid}`. ### Access levels Levels are ordered: `read < write < owner`. Each verb needs one: | Level | Verbs | |---|---| | Read | `sql`, `explain`, `explore`, `fields`, `tables`, branch status | | Write | `deploy`, `validate` (`zsql check`), `test`, remove a branch | | Owner | grant and revoke access, grant query keys, change settings, delete the project | Owners manage people with `GET /projects/{uid}/access` (anyone with Read may look), `PUT /projects/{uid}/access` with `{"email": "...", "level": "write"}`, and `DELETE /projects/{uid}/access/{email}`. The same list is the **People** tab in the console. ## Members and invites Admins add a member with `POST /account/members` and `{"email": "...", "name": "...", "role": "developer"}`. The response carries a one-time invite token (`zin_...`); send the person `https://app.0sql.io/invite?token=zin_...`. Accepting it sets their name and password and shows their first personal key once, the same way sign-up does. `POST /account/members/{email}/invite` issues a fresh token for a member who has not accepted yet; `DELETE /account/members/{email}` removes a member. `GET /account/members` lists members with their roles. ## Audit log Admins read `GET /account/audit?limit=50` (up to 1000) or the **Audit** tab: ```json {"events": [{"actor": "", "action": "deploy", "subject": "", "at": ""}]} ``` Each deploy writes one row. Planning does not. ## Playground Every project in the console has a playground at `/projects/{uid}/play?branch=`. Type a shorthand expression or a spec as JSON, add a context if the branch has policies, and press SQL, Explain or Explore (or Cmd-Enter). It has a field search and a cheat sheet, and remembers your last 20 queries in your browser only. It is a developer tool, not an end-user UI: it returns SQL and never runs it. ## What is stored Stored, per deployed branch: - the compiled model (tables, fields, joins, universes, policies) - the branch's tests - one audit row per deploy Never stored: - query specs and expressions - security contexts - the SQL that comes back `sql`, `explain` and `explore` write nothing: no log line, no database row, no file. Their only side effect is updating the key's `last_used_at`, at most once a minute. There is no request logging middleware in front of them. Never received at all: - the rows in your warehouse - the results of any statement - your warehouse credentials 0sql has no connection to your database. The compiled model describes your warehouse (names, SQL expressions, join conditions, policies) and that description is what the planner reads. Your `datasources.yml` can carry connection details as metadata for your own application, but `zsql deploy` strips every secret key from it before upload, and the service never opens a socket to a warehouse. Your data stays inside your network, and the statement you get back is executed there by you. ## Authentication errors Every error is `{"error": {"class": "...", "message": "..."}}`. | Status | Class | Message | Fix | |---|---|---|---| | 401 | `Unauthorized` | `an API key is required: Authorization: Bearer ` | Send the header. The service does not read `x-api-key`. | | 401 | `Unauthorized` | `this API key is not valid` | The secret is wrong, revoked or rotated. Create or copy a current one. | | 403 | `Forbidden` | `query keys are read only; deploy with your personal key` | You deployed, validated or tested with a `zqk_` key. | | 403 | `Forbidden` | `query key is not granted /` | Grant the key that project and branch, or `*`. | | 403 | `Forbidden` | ` has access to ; this needs ` | Ask an owner to raise your level. | | 403 | `Forbidden` | ` has no access to ` | The project is `restricted` and you are not on it. | | 403 | `Forbidden` | `deploying to the protected production branch '' of needs an owner` | Deploy to another branch, or have an owner deploy or turn `protected_production` off. | `zsql` prints these as `: `. ## Next steps - [Authentication with zsql](#authentication-and-configuration) - [Querying](#query-api) - [Security policies and the context](#row-level-security) - [API reference](#api-reference) --- ## Examples Each page in this cookbook takes one question you would ask of the TPC-DS model, writes it as a query spec (and as a `zsql sql --expr` line when the shorthand can say it), and shows the single SQL statement the planner returns. The SQL is quoted from the engine's own tests, unchanged. Where a variation has no test to quote, the page shows the request and describes the shape of the statement instead of inventing one. All requests go to `POST https://app.0sql.io/projects/tpcds/branches/main/sql` with `Authorization: Bearer zqk_...`. Fields are referenced by name or uid: `Web Net Paid` (`ws_net_paid`, on `web_sales`), `Net Paid` (`net_paid`, on `store_sales`), `Category`, `Product Name`, `Date`, `Customer Id`. | Page | Feature | What the SQL shows | |---|---|---| | [Cross-fact blend](#cross-fact-blend) | measures from two fact tables in one row | one aggregation CTE per fact, `FULL OUTER JOIN` on the conformed key, `COALESCE` | | [Complex measures](#complex-measures) | `calculations` with `[Name]@m`, `[Alias]`, CASE and windows; compound measures in YAML | arithmetic at the merge, re-aggregation over a subquery | | [Cohorts and segments](#cohorts-and-segments) | `segments` at query level and measure level, expanding pairs | a `seg…` CTE `INNER JOIN`ed on the key, inside one measure's CTE or around the whole query | | [Period over period](#period-over-period) | `temporalize` with `month_over_month`, `percent_change` | a shifted-date aggregation CTE, `HAVING` kept inside it | | [Share and windows](#share-and-windows) | `contribute`, `window` running and moving functions | a `CROSS JOIN` to a total CTE, `OVER (PARTITION BY … ORDER BY …)` | | [Filters and top-n](#filters-and-top-n) | flat lists, `and`/`or` trees, lists, dates, LIKE, measure filters, `top_n` | `WHERE` versus `HAVING`, a ranking CTE with `LIMIT` | | [Security context](#security-context) | `context` with groups and tags against `filter_data` and `mask_data` policies | `WHERE LOWER(dim) IN (…)`, `CASE WHEN … ELSE NULL END` | | [Level of detail](#level-of-detail) | exclusion and inclusion measures from the model, requested beside plain measures | an extra aggregation at a coarser or finer grain joined back | The pages reference each other but stand alone. If you have not seen a spec before, [The query spec](#the-query-spec) lists every key. Worked examples for both halves of 0sql: the query specs you POST and the YAML projects they plan against. Everything is built on the TPC-DS model. Every SQL statement in the request cookbook is the planner's own output for the request printed above it. ## Request cookbook Eight pages, one feature each. Each page states a problem, shows the request (JSON spec and the `zsql` shorthand where it can express it), the SQL that comes back, why the SQL has that shape, and the spec changes that give you the variations. - [Request cookbook](#request-cookbook): the hub, with a table of what each page demonstrates - [Cross-fact blend](#cross-fact-blend): two fact tables, one row per date - [Complex measures](#complex-measures): calculations in the spec and compound measures in the model - [Cohorts and segments](#cohorts-and-segments): populations joined into the query - [Period over period](#period-over-period): month-over-month and friends - [Share and windows](#share-and-windows): percent of total, running and moving windows - [Filters and top-n](#filters-and-top-n): every filter shape and the ranking CTE - [Security context](#security-context): the context your app sends and the WHERE and CASE it produces - [Level of detail](#level-of-detail): exclusion and inclusion measures requested like any other ## TPC-DS tutorial - [TPC-DS tutorial](#tpc-ds-tutorial): a complete walkthrough of the TPC-DS project, from `zsql init` to the first query ## Patterns - [Star schema](#star-schema-pattern): one fact, conformed dimensions - [Snowflake schema](#snowflake-schema-pattern): normalized dimension chains - [Fact-dimension](#fact-dimension-pattern): fact and dimension table patterns ## Recipes - [Customer 360](#customer-360-recipe): customer analytics across several facts - [Sales analysis](#sales-analysis-recipe): sales and revenue measures ## Suggested order 1. Read the [Request cookbook](#request-cookbook) to see what one request can ask for. 2. Follow the [TPC-DS tutorial](#tpc-ds-tutorial) to deploy the model those requests run against. 3. Use the patterns and recipes when you model your own warehouse. ### Fact-Dimension Pattern Core pattern for modeling fact and dimension tables. ## Overview The fact-dimension pattern is the foundation of dimensional modeling. Facts are events/transactions, dimensions are descriptive attributes. ## Fact Tables Facts represent business events or transactions. **Characteristics:** - Higher cost (100) - Many rows - Contain measures (aggregatable quantities) - Foreign keys to dimensions **Example:** ```yaml name: Orders cost: 100 fields: - type: measure name: Total Revenue expression: sql: sum(amount) ``` ## Dimension Tables Dimensions provide context for facts. **Characteristics:** - Lower cost (10) - Fewer rows - Contain dimensions (categorical fields) - Primary keys **Example:** ```yaml name: Customer cost: 10 fields: - type: dimension name: Customer ID expression: primary_key: true sql: customer_id ``` ## Relationships Facts join to dimensions: ```yaml orders_customer: left: Orders right: Customer sql: left.customer_id = right.id cardinality: many_to_one ``` ## Best Practices 1. **Fact cost: 100** - Higher cost 2. **Dimension cost: 10** - Lower cost 3. **Many-to-one** - Fact to dimension 4. **Primary keys** - On dimensions 5. **Measures in facts** - Aggregatable quantities ## Next Steps - [Learn star schema](#star-schema-pattern) - [Explore TPC-DS tutorial](#tpc-ds-tutorial) ### Snowflake Schema Pattern Model normalized data warehouses using the snowflake schema pattern. ## Overview Snowflake schema extends star schema by normalizing dimension tables. Dimensions can join to other dimensions. ## Structure ``` Fact Table ├── Dimension 1 │ └── Sub-Dimension 1 ├── Dimension 2 └── Dimension 3 └── Sub-Dimension 3 ``` ## Example ### Fact Table ```yaml name: Sales cost: 100 ``` ### Dimension to Dimension ```yaml datasource: warehouse # Fact to Dimension sales_customer: left: Sales right: Customer sql: left.customer_id = right.id cardinality: many_to_one # Dimension to Dimension customer_address: left: Customer right: Address sql: left.address_id = right.id cardinality: many_to_one ``` ## Use Cases - **Normalized dimensions** - Reduce data duplication - **Hierarchical data** - Product categories, geography - **Complex relationships** - Multiple levels of dimensions ## Best Practices 1. **Use when needed** - Only normalize if it adds value 2. **Set appropriate costs** - All dimensions: cost 10 3. **Document relationships** - Explain dimension hierarchies ## Next Steps - [Learn star schema](#star-schema-pattern) - [Explore fact-dimension pattern](#fact-dimension-pattern) ### Star Schema Pattern Model data warehouses using the star schema pattern. ## Overview Star schema is a common data warehouse pattern with a central fact table surrounded by dimension tables. ## Structure ``` Fact Table (center) ├── Dimension 1 ├── Dimension 2 ├── Dimension 3 └── Dimension N ``` ## Example ### Fact Table ```yaml name: Sales physical_name: sales datasource: warehouse cost: 100 # Fact table: higher cost fields: - type: dimension name: Sale ID data_type: integer expression: primary_key: true sql: sale_id - type: dimension name: Sale Date data_type: date expression: sql: sale_date - type: measure name: Total Revenue data_type: decimal expression: sql: sum(amount) ``` ### Dimension Tables ```yaml # Customer Dimension name: Customer cost: 10 # Dimension: lower cost # Product Dimension name: Product cost: 10 # Date Dimension name: Date cost: 10 ``` ### Relationships ```yaml datasource: warehouse # Fact to Dimensions sales_customer: left: Sales right: Customer sql: left.customer_id = right.id cardinality: many_to_one sales_product: left: Sales right: Product sql: left.product_id = right.id cardinality: many_to_one sales_date: left: Sales right: Date sql: left.sale_date_id = right.date_id cardinality: many_to_one ``` ## Best Practices 1. **Fact table cost: 100** - Higher cost 2. **Dimension table cost: 10** - Lower cost 3. **Many-to-one relationships** - Fact to dimensions 4. **Primary keys on dimensions** - Unique identifiers 5. **Organize by domain** - Group related tables ## Next Steps - [Learn snowflake schema](#snowflake-schema-pattern) - [Explore fact-dimension pattern](#fact-dimension-pattern) ### Customer 360 Recipe Build a comprehensive customer analytics model with multiple touchpoints. ## Overview Customer 360 provides a unified view of customer behavior across all touchpoints: orders, support interactions, marketing engagement, and lifetime value. This recipe demonstrates a production-ready model structure. ## Architecture ```mermaid flowchart TD O[Orders Fact] -->|many_to_one| C[Customer] S[Support Tickets] -->|many_to_one| C E[Email Events] -->|many_to_one| C C -->|many_to_one| CA[Customer Address] C -->|many_to_one| CS[Customer Segment] ``` ## Complete Model Structure ### Customer Dimension ```yaml name: Customer physical_name: customers datasource: warehouse cost: 10 fields: - type: dimension name: Customer ID data_type: string expression: primary_key: true lookup: true sql: customer_id - type: dimension name: Customer Name data_type: string expression: lookup: true sql: concat(first_name, ' ', last_name) - type: dimension name: Customer Email data_type: string expression: lookup: true sql: email - type: dimension name: Customer Since data_type: date grains: [day, month, quarter, year] expression: sql: created_at - type: dimension name: Customer Status data_type: string expression: lookup: true sql: status - type: dimension name: Customer Tier description: Based on lifetime spend data_type: string expression: lookup: true sql: | CASE WHEN lifetime_value >= 10000 THEN 'Platinum' WHEN lifetime_value >= 5000 THEN 'Gold' WHEN lifetime_value >= 1000 THEN 'Silver' ELSE 'Bronze' END ``` ### Orders Fact Table ```yaml name: Orders physical_name: orders datasource: warehouse cost: 100 fields: - type: dimension name: Order ID data_type: string expression: primary_key: true sql: order_id - type: dimension name: Order Date data_type: date grains: [day, week, month, quarter, year] expression: sql: order_date - type: dimension name: Order Channel data_type: string expression: lookup: true sql: channel - type: measure name: Total Revenue data_type: decimal format: currency:2 expression: sql: sum(order_total) - type: measure name: Order Count data_type: integer expression: sql: count(distinct order_id) - type: measure name: Average Order Value description: Revenue per order data_type: decimal format: currency:2 expression: sql: "[Total Revenue] / nullif([Order Count], 0)" - type: measure name: Items Per Order data_type: decimal expression: sql: sum(item_count) / nullif(count(distinct order_id), 0) ``` ### Support Tickets Fact ```yaml name: Support Tickets physical_name: support_tickets datasource: warehouse cost: 100 fields: - type: dimension name: Ticket ID data_type: string expression: primary_key: true sql: ticket_id - type: dimension name: Ticket Created Date data_type: date grains: [day, week, month] expression: sql: created_at - type: dimension name: Ticket Category data_type: string expression: lookup: true sql: category - type: dimension name: Ticket Priority data_type: string expression: lookup: true sql: priority - type: measure name: Ticket Count data_type: integer expression: sql: count(distinct ticket_id) - type: measure name: Avg Resolution Hours data_type: decimal expression: sql: avg(resolution_hours) - type: measure name: First Response Hours data_type: decimal expression: sql: avg(first_response_hours) ``` ### Relationships ```yaml datasource: warehouse # Orders to Customer orders_customer: left: Orders right: Customer sql: left.customer_id = right.customer_id cardinality: many_to_one # Support Tickets to Customer tickets_customer: left: Support Tickets right: Customer sql: left.customer_id = right.customer_id cardinality: many_to_one # Customer to Address (for geographic analysis) customer_address: left: Customer right: Customer Address sql: left.address_id = right.address_id cardinality: many_to_one allow_measure_expansion: true ``` ## Key Measures Explained | Measure | Definition | Business Use | |--------|------------|--------------| | Total Revenue | `sum(order_total)` | Track overall sales | | Order Count | `count(distinct order_id)` | Volume analysis | | Average Order Value | `Revenue / Orders` | Basket size tracking | | Customer Lifetime Value | `Total Revenue` grouped by Customer ID | Segment customers | | Ticket Count | Support interactions | Support load | | Avg Resolution Hours | Time to resolve | Support efficiency | ## Advanced: Compound Measures Lifetime value needs no new measure: project `Total Revenue` by `Customer ID` and the planner joins Orders to Customer. Compound measures are for formulas over other fields: their expression references measures and dimensions in brackets rather than warehouse columns. A formula built only from measure references needs no aggregate of its own, because each referenced measure carries one. Anything that needs a raw column goes in a plain measure first. ```yaml # In Orders: a plain measure over the fact's own columns - type: measure name: Active Days description: Days between a customer's first and last order data_type: integer expression: sql: datediff(day, min(order_date), max(order_date)) # Compound: other measures, referenced by name - type: measure name: Customer Order Frequency description: Average days between orders data_type: decimal expression: sql: "[Active Days] / nullif([Order Count] - 1, 0)" # Compound across two facts: Support Tickets and Orders blend on Customer - type: measure name: Tickets per 1000 Orders description: Support load relative to order volume data_type: decimal expression: sql: "[Ticket Count] * 1000.0 / nullif([Order Count], 0)" ``` ## Query Examples Each request is one line of shorthand or a JSON spec; 0sql returns the SQL and your application runs it. ```sh zsql sql --expr "customer tier, total revenue, order count, average order value" zsql sql --expr "customer name, customer tier, ticket count, avg resolution hours" zsql sql --expr "month(customer since), total revenue, order count" ``` The third line, a monthly cohort view, as the spec your application would POST: ```json { "spec": { "projections": [ {"field": "Customer Since", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Total Revenue"}, {"field": "Order Count"} ] } } ``` Add `{"type": "temporalize", "transform": "year_over_year"}` to the `Total Revenue` projection's `decorators` for year-over-year growth per cohort month. See [Projections and decorators](#projections-and-decorators). ## Best Practices 1. **Use Customer ID as primary grain**: All customer-level analysis flows from here 2. **Add lookup: true** on frequently filtered dimensions 3. **Use compound measures** for cross-table calculations 4. **Include date dimensions with grains** for time-series analysis 5. **Document business logic** in descriptions ## Next Steps - [Compound Measures](#compound-measures): Cross-table calculations - [Extended Blending](#extended-blending-groups): one customer concept across fact tables - [Projections and decorators](#projections-and-decorators): YoY customer growth and other request-time transforms ### Sales Analysis Recipe Build a production-ready sales analytics model with multi-channel support. ## Overview This recipe creates a comprehensive sales model supporting: - Multi-channel sales (store, web, catalog) - Product hierarchy analysis - Geographic breakdown - Time-series comparisons - Profitability measures ## Architecture ```mermaid flowchart TD SS[Store Sales] -->|many_to_one| P[Product] SS -->|many_to_one| S[Store] SS -->|many_to_one| D[Date] SS -->|many_to_one| C[Customer] WS[Web Sales] -->|many_to_one| P WS -->|many_to_one| D WS -->|many_to_one| C P -->|many_to_one| PC[Product Category] S -->|many_to_one| R[Region] ``` ## Complete Model Structure ### Sales Fact Table ```yaml name: Sales physical_name: sales_fact datasource: warehouse cost: 100 fields: # Keys and IDs - type: dimension name: Transaction ID data_type: string expression: primary_key: true sql: transaction_id - type: dimension name: Sale Date data_type: date grains: [day, week, month, quarter, year] expression: sql: sale_date - type: dimension name: Sale Channel data_type: string expression: lookup: true sql: channel # Revenue measures - type: measure name: Gross Revenue description: Total sales before returns and discounts data_type: decimal format: currency:2 expression: sql: sum(gross_amount) - type: measure name: Net Revenue description: Sales after returns and discounts data_type: decimal format: currency:2 expression: sql: sum(net_amount) - type: measure name: Discount Amount data_type: decimal format: currency:2 expression: sql: sum(discount_amount) # Volume measures - type: measure name: Units Sold data_type: integer expression: sql: sum(quantity) - type: measure name: Transaction Count data_type: integer expression: sql: count(distinct transaction_id) # Profitability - type: measure name: Cost of Goods Sold data_type: decimal format: currency:2 expression: sql: sum(cogs) - type: measure name: Gross Profit data_type: decimal format: currency:2 expression: sql: "[Net Revenue] - [Cost of Goods Sold]" - type: measure name: Gross Margin Percent data_type: decimal format: percent:2 expression: sql: "[Gross Profit] / nullif([Net Revenue], 0)" # Derived measures - type: measure name: Average Transaction Value data_type: decimal format: currency:2 expression: sql: "[Net Revenue] / nullif([Transaction Count], 0)" - type: measure name: Average Selling Price data_type: decimal format: currency:2 expression: sql: "[Net Revenue] / nullif([Units Sold], 0)" - type: measure name: Discount Rate data_type: decimal format: percent:2 expression: sql: "[Discount Amount] / nullif([Gross Revenue], 0)" ``` ### Product Dimension ```yaml name: Product physical_name: products datasource: warehouse cost: 10 fields: - type: dimension name: Product ID data_type: string expression: primary_key: true lookup: true sql: product_id - type: dimension name: Product Name data_type: string expression: lookup: true sql: product_name - type: dimension name: Product Category data_type: string expression: lookup: true sql: category - type: dimension name: Product Subcategory data_type: string expression: lookup: true sql: subcategory - type: dimension name: Brand data_type: string expression: lookup: true sql: brand - type: dimension name: Product Status description: Active, Discontinued, etc. data_type: string expression: lookup: true sql: status ``` ### Store Dimension ```yaml name: Store physical_name: stores datasource: warehouse cost: 10 fields: - type: dimension name: Store ID data_type: string expression: primary_key: true lookup: true sql: store_id - type: dimension name: Store Name data_type: string expression: lookup: true sql: store_name - type: dimension name: Store Type data_type: string expression: lookup: true sql: store_type - type: dimension name: Store City data_type: string expression: lookup: true sql: city - type: dimension name: Store State data_type: string expression: lookup: true sql: state - type: dimension name: Store Region data_type: string expression: lookup: true sql: region - type: dimension name: Store Open Date data_type: date expression: sql: open_date ``` ### Date Dimension ```yaml name: Date physical_name: date_dim datasource: warehouse cost: 10 fields: - type: dimension name: Date data_type: date grains: [day, week, month, quarter, year] expression: primary_key: true sql: date_value - type: dimension name: Day of Week data_type: string expression: lookup: true sql: day_name - type: dimension name: Month Name data_type: string expression: lookup: true sql: month_name - type: dimension name: Quarter data_type: string expression: lookup: true sql: quarter_name - type: dimension name: Year data_type: integer expression: lookup: true sql: year_number - type: dimension name: Is Weekend data_type: boolean expression: sql: is_weekend - type: dimension name: Is Holiday data_type: boolean expression: sql: is_holiday ``` ### Relationships ```yaml datasource: warehouse # Sales to Product sales_product: left: Sales right: Product sql: left.product_id = right.product_id cardinality: many_to_one # Sales to Store sales_store: left: Sales right: Store sql: left.store_id = right.store_id cardinality: many_to_one # Sales to Date sales_date: left: Sales right: Date sql: left.sale_date = right.date_value cardinality: many_to_one # Sales to Customer sales_customer: left: Sales right: Customer sql: left.customer_id = right.customer_id cardinality: many_to_one join: left # Not all sales have customer (e.g., anonymous) ``` ## Key Measures Reference | Measure | Formula | Use Case | |--------|---------|----------| | Gross Revenue | `sum(gross_amount)` | Total sales volume | | Net Revenue | `sum(net_amount)` | Actual revenue | | Gross Profit | `Net Revenue - COGS` | Profitability | | Gross Margin % | `Gross Profit / Net Revenue` | Margin analysis | | ATV | `Net Revenue / Transactions` | Basket size | | ASP | `Net Revenue / Units Sold` | Pricing analysis | | Discount Rate | `Discounts / Gross Revenue` | Promotion impact | ## Time-Series Analysis Period comparisons and windows are not modeled. They are `decorators` on a projection in the spec your application sends; the planner adds the comparison as a final pass over the aggregated result. See [Projections and decorators](#projections-and-decorators). ```json { "spec": { "projections": [ {"field": "Sale Date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Net Revenue"}, {"field": "Net Revenue", "alias": "Net Revenue LY", "decorators": [{"type": "temporalize", "transform": "year_over_year"}]}, {"field": "Transaction Count", "alias": "Transactions MoM %", "decorators": [{"type": "temporalize", "transform": "month_over_month", "percent_change": true}]} ] } } ``` At day grain, a moving average is `{"type": "window", "mode": "moving", "function": "avg", "size": 7}` on the measure. The same two requests as shorthand: ```sh zsql sql --expr "month(sale date), net revenue, yoy(net revenue) as Net Revenue LY, mom_pct(transaction count)" zsql sql --expr "day(sale date), net revenue, moving_avg(net revenue, 7)" ``` ## Common Queries ```sh zsql sql --expr "sale channel, month(sale date), net revenue, transaction count, average transaction value" zsql sql --expr "product category, brand, net revenue, units sold, gross margin percent" zsql sql --expr "store name, store region, net revenue, transaction count, discount rate" zsql sql --expr "day of week, net revenue, transaction count, is weekend = true" ``` The last one, day-of-week analysis with a filter, as a spec: ```json { "spec": { "projections": [ {"field": "Day of Week"}, {"field": "Net Revenue"}, {"field": "Transaction Count"} ], "filters": [ {"field": "Is Weekend", "predicate": "equals", "value": "true"} ] } } ``` ## Best Practices 1. **Separate gross and net measures**: Track discounts and returns explicitly 2. **Use compound measures** for ratios: Prevents division errors 3. **Include cost data**: Enable profitability analysis 4. **Add date grains**: Support flexible time grouping 5. **Use LEFT join for optional dimensions**: Handle anonymous sales ## Next Steps - [Projections and decorators](#projections-and-decorators): YoY and MoM comparisons, running totals, moving averages, percent of total - [Query spec](#the-query-spec): the full request shape - [Filters](#filters): predicates and relative dates ### Cohorts and segments A segment is a population: a set of key-dimension values picked out by filters, by measures, or both. The planner materializes it as a CTE and joins it into the query on the key. Applied to the whole query it behaves like a cohort filter. Applied to one measure through `apply_to` it lets a cohort's number sit beside the unconstrained baseline in the same row. An expanding segment goes the other way and carries a dimension of the *other* members sharing the key into the query, which is how basket pairs are built. Reference: [Segments](#segments). The shorthand has no segment syntax; these requests are JSON only. The statements on this page were planned with a system-admin context that bypasses the branch's policies; the SQL is the same on a branch without policies. ## Query-level include Web revenue per product, restricted to products that have at least one row in the `books` category. ```json {"spec": { "name": "Product Sales", "projections": [{"field": "product_name", "alias": "Product Name"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}], "segments": [{"name": "Book Products", "mode": "include", "keys": ["product_name"], "filters": [{"field": "category", "predicate": "equals", "value": "books"}]}] }} ``` ```sql WITH seg4a9429dd0d8da260126c0bdb19cc5c6a AS ( SELECT T0."i_product_name" AS "dimbe52306" FROM item T0 WHERE LOWER(T0."i_category") = 'books' GROUP BY T0."i_product_name" ) SELECT T1."i_product_name" AS "Product Name", sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN seg4a9429dd0d8da260126c0bdb19cc5c6a AS Q2 ON T1."i_product_name" = Q2.dimbe52306 GROUP BY T1."i_product_name" ``` - **The segment is its own CTE.** `seg4a94…` selects the key (`i_product_name`) from the tables needed to evaluate the segment's filters, grouped by the key so each member appears once. - **The query is INNER JOINed to it on the key.** `apply_to` is empty, so the join sits in the main query and every projected measure is constrained. - **The segment's filter does not leak into the query's WHERE.** The query asks nothing about category; membership is decided inside the CTE. ## Measure-level, inside a blend Same projections plus the store measure, and the segment now applies to `ws_net_paid` only. The row shows book-product web revenue beside all-product store revenue. ```json {"spec": { "name": "Product Sales", "projections": [ {"field": "product_name", "alias": "Product Name"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "net_paid", "alias": "Net Paid"} ], "segments": [{"name": "Book Products", "mode": "include", "keys": ["product_name"], "filters": [{"field": "category", "predicate": "equals", "value": "books"}], "apply_to": ["ws_net_paid"]}] }} ``` ```sql WITH ag6610e03993b3da57101097f16f2df44d AS ( SELECT T1."i_product_name" AS "dimbe52306", sum(T0."ss_net_paid") AS "msr621f67c" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk GROUP BY T1."i_product_name" ), seg4a9429dd0d8da260126c0bdb19cc5c6a AS ( SELECT T0."i_product_name" AS "dimbe52306" FROM item T0 WHERE LOWER(T0."i_category") = 'books' GROUP BY T0."i_product_name" ), ag4e173d8139dc2fe7bf4cf4d6aced1d58 AS ( SELECT T1."i_product_name" AS "dimbe52306", sum(T0."ws_net_paid") AS "msr0c0d023" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN seg4a9429dd0d8da260126c0bdb19cc5c6a AS Q2 ON T1."i_product_name" = Q2.dimbe52306 GROUP BY T1."i_product_name" ) SELECT COALESCE(A1.dimbe52306, A0.dimbe52306) AS "Product Name", A0.msr0c0d023 AS "Web Net Paid", A1.msr621f67c AS "Net Paid" FROM ag4e173d8139dc2fe7bf4cf4d6aced1d58 A0 FULL OUTER JOIN ag6610e03993b3da57101097f16f2df44d A1 ON A1.dimbe52306 = A0.dimbe52306 ``` - **The join moved inside one aggregation CTE.** Only `ag4e17…` (web) joins `seg4a94…`. The store CTE `ag6610…` aggregates every product. - **The blend is unchanged.** The two aggregation CTEs are stitched with the same `FULL OUTER JOIN` and `COALESCE` as a plain [cross-fact blend](#cross-fact-blend). Products without book rows still appear with `Web Net Paid` as `NULL`. - **`apply_to` matches by alias first, then by measure uid.** `"apply_to": ["Web Net Paid"]` would do the same. If two projections share a measure, give one an alias and name the alias. ## Expanding segment: basket pairs For each product, the other products the same customers bought in store, with web revenue of the first product. The key is `customer_id`; the segment carries `product_name` as a new column. ```json {"spec": { "projections": [{"field": "product_name", "alias": "Product Name"}, {"field": "ws_net_paid", "alias": "Web Net Paid"}], "segments": [{"name": "Customer Products", "keys": ["customer_id"], "expanding": [{"field": "product_name", "as": "Product Name (same Customer ID)"}]}] }} ``` ```sql WITH seg87ae4658940c714d42c8e2714d387839 AS ( SELECT T1."c_customer_id" AS "dimf5dcb79", T2."i_product_name" AS "dimd3303bd" FROM store_sales T0 JOIN customer T1 ON T0.ss_customer_sk = T1.c_customer_sk JOIN item T2 ON T0.ss_item_sk = T2.i_item_sk GROUP BY T1."c_customer_id", T2."i_product_name" ) SELECT T1."i_product_name" AS "Product Name", sum(T0."ws_net_paid") AS "Web Net Paid", Q2.dimd3303bd AS "Product Name (same Customer ID)" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN customer T3 ON T0.ws_bill_customer_sk = T3.c_customer_sk INNER JOIN seg87ae4658940c714d42c8e2714d387839 AS Q2 ON T3."c_customer_id" = Q2.dimf5dcb79 AND T1."i_product_name" <> Q2.dimd3303bd GROUP BY T1."i_product_name", Q2.dimd3303bd ``` - **The CTE holds (key, carried dimension) pairs.** `seg87ae…` groups `store_sales` by customer and product. - **The join matches on the key and excludes self-pairs.** `pair_dedupe` defaults to `neq`, which is the `<>` in the join condition. `lt` keeps one canonical ordering of each pair; `none` keeps self-pairs. - **The carried column is projected and grouped.** `Q2.dimd3303bd AS "Product Name (same Customer ID)"` comes from the segment, not from the model, and joins the `GROUP BY`. Omit `as` and the column is named `Product Name (same Customer Id)` by default. - **`join` picks inner or left.** `"join": "left"` keeps products whose customers bought nothing else. `mode` and `apply_to` are not allowed on an expanding segment. ## Variations - **Exclude mode.** `"mode": "exclude"` on the first request turns the population into an anti-join: the segment CTE is unchanged and the query keeps products that are *not* in it. No captured statement to quote here; the CTE body is identical to the include case. - **Membership defined by a measure.** Products that sold more than 1000 in store, applied to web revenue: ```json "segments": [{"name": "Big in Store", "keys": ["product_name"], "measures": ["net_paid"], "filters": [{"field": "net_paid", "predicate": "greater_than", "value": "1000"}], "apply_to": ["ws_net_paid"]}] ``` Listing `net_paid` under `measures` routes the segment through `store_sales`; the measure filter becomes a `HAVING sum(ss_net_paid) > 1000` on the segment CTE, which groups by the key. Several measures from several facts (bought in store OR catalog) go in the same `measures` list; that is the supported way to say "any of these facts", since `or` groups are not allowed inside a flat filter list. - **Rank the population.** A `top_n` filter inside a segment ranks the members: `{"field": "product_name", "predicate": "top_n", "value": "20", "top_n_measure": "net_paid"}` with `"measures": ["net_paid"]` keeps the twenty best-selling store products as the cohort. ## Next steps - [Segments](#segments) for every key and error message - [Filters and top-n](#filters-and-top-n) for the filter shapes a segment accepts - [Cross-fact blend](#cross-fact-blend) ### Complex measures A calculation is a SQL formula over fields and other projections, written in the spec and planned with the query. References go through brackets: `[Name]@m` is a measure, `[Name]@d` a dimension, `[Alias]` another projection or calculation in the same request. Raw column names are refused. This page builds up from a plain ratio to a windowed calculation that references another calculation, shows the SQL the planner returns for two of them, and then contrasts request-side calculations with compound measures defined in YAML. The formula grammar is on [Calculations](#calculations). ## Seven formulas Each entry is a complete projection you can drop into `projections` (or into the top-level `calculations` list, where `calculation: true` is implied). `data_type` defaults to `decimal`. 1. **Ratio of two measures from different facts.** Each side is aggregated in its own CTE; the division happens in the final `SELECT`. ```json {"alias": "2x WPaid", "sql": "[Net Paid]@m/[Web Net Paid]@m", "data_type": "decimal", "calculation": true} ``` 2. **Aggregated ratio.** A formula that calls an aggregate function (`sum`, `max`, `count`, …) is a measure calculation and re-aggregates its inputs. ```json {"alias": "Agg Ratio", "sql": "sum([Web Net Paid]@m)/max([Net Paid]@m)", "calculation": true} ``` 3. **Conditional distinct count.** `CASE` is the conditional; there is no `if()`. ```json {"alias": "Actives", "data_type": "integer", "sql": "count(distinct case when [Net Paid]@m > 0 or [Web Net Paid]@m > 0 then [Category]@d else null end)", "calculation": true} ``` 4. **A calculation referencing a calculation, with a window.** `[Actives]` is the alias of entry 3 (or of the simpler `count(distinct [Category]@d)` used in the statement below). ```json {"alias": "Retention Rate", "sql": "[Actives] / nullif(first_value([Actives]) over (order by [Category]@d), 0)", "calculation": true} ``` 5. **String dimension calculation.** No aggregate in the formula, so it groups like a dimension. ```json {"alias": "Cat Concat Item", "data_type": "string", "sql": "[Category] || ' ' || [Product Name]@d", "calculation": true} ``` 6. **Date part.** Also a dimension calculation. ```json {"alias": "Year", "data_type": "integer", "sql": "EXTRACT(year FROM [Date]@d)", "calculation": true} ``` 7. **Margin percent.** Guard the denominator with `nullif`. ```json {"alias": "Margin %", "calculation": true, "sql": "([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)", "data_type": "decimal"} ``` In the shorthand a calculation is `Alias = formula` or `Alias := formula`: ```sh zsql sql --expr "category, net paid, Bucket = [Net Paid] / 100, web net paid, category in (men, children)" zsql sql --expr "Margin % := ([Net Paid]@m - [Cost]@m) / nullif([Net Paid]@m, 0)" ``` ## A calculation on a blended measure The request is the [cross-fact blend](#cross-fact-blend) plus a calculation over the store measure's alias. ```json {"spec": { "projections": [ {"field": "category", "alias": "Category"}, {"field": "net_paid", "alias": "Net Paid"}, {"alias": "Bucket", "calculation": true, "sql": "[Net Paid] / 100", "data_type": "decimal"}, {"field": "ws_net_paid", "alias": "Web Net Paid"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] }} ``` ```sql WITH agc51c5a013f468165df0d33fc12150511 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ss_net_paid") AS "msr8a51bb0" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ), ag89dc4ad5879fa4552696fe3de5b10435 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ws_net_paid") AS "msr60b3792" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ) SELECT COALESCE(A1.dim30d09b7, A0.dim30d09b7) AS "Category", A1.msr8a51bb0 AS "Net Paid", A1.msr8a51bb0 / 100 AS "Bucket", A0.msr60b3792 AS "Web Net Paid" FROM ag89dc4ad5879fa4552696fe3de5b10435 A0 FULL OUTER JOIN agc51c5a013f468165df0d33fc12150511 A1 ON A1.dim30d09b7 = A0.dim30d09b7 ``` The calculation stays at the merge: `A1.msr8a51bb0 / 100` is computed in the final `SELECT` from the store CTE's already aggregated column. Neither aggregation CTE knows the calculation exists. A ratio across the two facts (entry 1) lands in the same place, as `A1.msr… / A0.msr…`. ## A calculation over a calculation, with a window Projections: `category`, `Actives = count(distinct [Category]@d)` (integer) and `Retention Rate = [Actives] / nullif(first_value([Actives]) over (order by [Category]@d), 0)`. ```json {"spec": { "projections": [ {"field": "category", "alias": "Category"}, {"alias": "Actives", "data_type": "integer", "sql": "count(distinct [Category]@d)", "calculation": true}, {"alias": "Retention Rate", "sql": "[Actives] / nullif(first_value([Actives]) over (order by [Category]@d), 0)", "calculation": true} ] }} ``` ```sql SELECT A0.dim30d09b7 AS "Category", count(distinct A0.dim30d09b7) AS "Actives", count(distinct A0.dim30d09b7) / nullif(first_value(count(distinct A0.dim30d09b7)) over (order by A0.dim30d09b7), 0) AS "Retention Rate" FROM ( SELECT T0."i_category" AS "dim30d09b7" FROM web_sales_sum T0 GROUP BY T0."i_category" ) A0 GROUP BY A0.dim30d09b7 ``` Two things to read off this statement. The planner resolved `Category` through a table named `web_sales_sum` in the test model, grouped it in a subquery, and applied the aggregating calculation in the outer `SELECT` with its own `GROUP BY`, so a measure calculation never aggregates raw rows twice. And `[Actives]` is expanded in place: the window function wraps the same `count(distinct …)` expression, and the window clause stays out of the `GROUP BY`. ## Model-side compound measures versus request-side calculations The same bracket grammar works in a measure's `expression` in the model. These two are verbatim from the TPC-DS project's `tbl.store_sales.yml`: ```yaml - type: measure name: Sales Throughput description: Rate at which inventory is turned over data_type: decimal expression: "[Store Quantity] / [Inventory Quantity On Hand] * 1.0" - type: measure name: Store Net Including Returns description: Net amount paid for sales minus return amount data_type: decimal expression: "[Store Net Paid] - [Store Return Amount]" ``` A measure made only of bracket references needs no aggregate function; the referenced measures carry their own. In a model expression write `[Name]@m` and `[Name]@d` when you want to be explicit about the kind. Once deployed, a compound measure is requested like any other field: `{"field": "Store Net Including Returns"}`. See [Compound measures](#compound-measures). Put a formula in the **model** when: - more than one caller should get the same number (margin, net of returns, throughput); - it should be discoverable through `GET …/fields` and the explore endpoint; - a security policy should trigger on it (policies fire on projected fields, never on calculations); - it should get a `format`, `description` and `synonyms`. Put a formula in the **request** when: - it references another projection by `[Alias]`, or a window over the query's own ordering (entry 4), which only exist at request time; - it is specific to one screen or one API call; - it is a dimension calculation that shapes the grouping of this query only (entries 5 and 6). ## Variations - **Ratio across facts.** Replace the `Bucket` projection with entry 1. The two aggregation CTEs are unchanged; the final `SELECT` divides the store column by the web column. - **Share of a total.** `Share = sum([Web Net Paid]@m) / sum([Net Paid]@m)` as the only projection beside `category` gives one ratio per category. For percent of the grand total use the `contribute` decorator instead; see [Share and windows](#share-and-windows). - **Order by the calculation.** Add `"order_by": "desc"` to the calculation projection. There is no top-level sort key. ## Next steps - [Calculations](#calculations) for the full grammar and error messages - [Compound measures](#compound-measures) - [Level of detail](#level-of-detail) when a ratio needs a denominator at a different grain ### Cross-fact blend `web_sales` and `store_sales` are different fact tables with different row grains. You want web revenue and store revenue side by side, one row per sold date, filtered to a few categories. Joining the two facts row to row would multiply every web row by every store row on the same date. The planner never does that: it aggregates each fact on its own to the common grain and stitches the results together afterwards. This is drill-across, and you get it by listing the two measures in one spec. ## The model side Each fact carries its own sold-date dimension, and both belong to one extended blending group, so a query grouped by `Web Sold Date` can be answered by `store_sales` through its own date member. The fragment below is the shape of the model these requests plan against. ```yaml # models/web/tbl.web_sales.yml - type: dimension name: Web Sold Date data_type: integer extended_blend_group: blendable_fact_dates expression: sql: ws_sold_date_sk - type: measure name: Web Net Paid data_type: decimal expression: sql: sum(ws_net_paid) # models/store/tbl.store_sales.yml - type: dimension name: Store Sold Date data_type: integer extended_blend_group: blendable_fact_dates expression: sql: ss_sold_date_sk - type: measure name: Net Paid data_type: decimal expression: sql: sum(ss_net_paid) ``` See [Extended blending groups](#extended-blending-groups) for how formation turns the group into blend paths. A dimension on a shared table (`Date` on `date_dim`, `Category` on `item`) needs no group: both facts join to it already. ## The request **curl:** ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"spec": { "name": "Query One", "projections": [ {"field": "ws_net_paid", "alias": "Web Net Paid"}, {"field": "ws_sold_date", "alias": "Web Sold Date"}, {"field": "net_paid", "alias": "Net Paid"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] }}' ``` **zsql:** ```sh zsql sql --expr "ws_net_paid as Web Net Paid, ws_sold_date as Web Sold Date, net_paid as Net Paid, category in (men, children)" ``` The shorthand strips quotes from list items, so the third item `"books,com"` cannot be written on one line. Use the JSON spec when a list value contains a comma. See [Filters and top-n](#filters-and-top-n). ## The SQL ```sql WITH ag7098b0d0901f2eb16d14f9356f0bb2a0 AS ( SELECT T0."ss_sold_date_sk" AS "dimed56b67", sum(T0."ss_net_paid") AS "msr621f67c" FROM store_sales T0 JOIN item T1 ON T0.ss_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T0."ss_sold_date_sk" ), ag29130fae5cf6548d6ddb1bd37298b407 AS ( SELECT sum(T0."ws_net_paid") AS "msr501e4a8", T0."ws_sold_date_sk" AS "dimed56b67" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T0."ws_sold_date_sk" ) SELECT A0.msr501e4a8 AS "Web Net Paid", COALESCE(A1.dimed56b67, A0.dimed56b67) AS "Web Sold Date", A1.msr621f67c AS "Net Paid" FROM ag29130fae5cf6548d6ddb1bd37298b407 A0 FULL OUTER JOIN ag7098b0d0901f2eb16d14f9356f0bb2a0 A1 ON A1.dimed56b67 = A0.dimed56b67 ``` ## Why the SQL looks like this - **One aggregation CTE per fact.** `ag7098…` sums `store_sales`, `ag2913…` sums `web_sales`. Each groups by its own sold-date key, which the blend group maps to the same output column (`dimed56b67`). Each CTE applies the category filter through its own join to `item`, so the filter is honoured on both sides. - **Grain safety.** Neither CTE sees the other fact. There is no row-level join between `web_sales` and `store_sales`, so no measure is inflated. - **FULL OUTER JOIN on the conformed key.** The final `SELECT` joins the two CTEs on the date key. A date with web sales but no store sales still appears, with `Net Paid` as `NULL`, and the other way round. - **COALESCE on the key.** `COALESCE(A1.dimed56b67, A0.dimed56b67)` picks the date from whichever side has it, so the projected dimension is never `NULL` just because one fact is missing that date. - **The join type is a dialect setting.** `final_pass_measure_join_type` is `full` in the base dialect. Override it per request through `db_settings` (next section). ## Variations - **A calculation standing on one of the measures.** Add `{"alias": "Bucket", "calculation": true, "sql": "[Net Paid] / 100", "data_type": "decimal"}` to the projections. The CTEs do not change; the final `SELECT` gains `A1.msr8a51bb0 / 100 AS "Bucket"`. The full request and SQL are the first example on [Complex measures](#complex-measures). - **Keep only dates the web fact has.** Add `"db_settings": {"final_pass_measure_join_type": "left"}` to the spec. The final join becomes a `LEFT JOIN` from the first aggregation CTE; dates present only in `store_sales` drop out. The CTEs are unchanged. - **Blend on a shared dimension instead.** Replace `ws_sold_date` with `{"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}`. Both CTEs then join to `date_dim` and group by `DATE_TRUNC('month', d_date)`; the final join key is the month. No blend group is needed because `Date` lives on a table both facts reach. ## Next steps - [Complex measures](#complex-measures) for ratios across the two facts - [Extended blending groups](#extended-blending-groups) - [The query spec](#the-query-spec) for `db_settings` and the other top-level keys ### Filters and top-n Filters are plain values against named fields. A filter on a dimension becomes a `WHERE`; a filter on a measure becomes a `HAVING` on the node that groups. Lists are one comma-separated string, strings compare lower-cased, dates accept fixed formats and relative offsets, and `top_n` keeps the best N members of a dimension by a measure through a ranking CTE. Reference: [Filters](#filters). ## Flat AND list An array is an AND of leaves. Every entry names a `field` and a `predicate`. ```json "filters": [ {"field": "category", "predicate": "in_list", "value": "men, children"}, {"field": "date", "predicate": "greater_than_or_equal_to", "value": "2024-01-01"}, {"field": "net_paid", "predicate": "greater_than", "value": "100"} ] ``` ```sh zsql sql --expr "category, net paid, category in (men, children), date >= 2024-01-01, net paid > 100" ``` Every filter in the shorthand is an item in the same comma list as the projections, and the shorthand only produces flat AND lists. ## The and/or tree An object with an `and` or `or` key holds nested nodes. This is JSON only. ```json "filters": {"or": [ {"field": "category", "predicate": "equals", "value": "books"}, {"and": [ {"field": "category", "predicate": "equals", "value": "music"}, {"field": "date", "predicate": "greater_than_or_equal_to", "value": "2024-01-01"} ]} ]} ``` An `and`/`or` object *inside* a flat array is rejected: `Filters here are a flat AND list ...`. The error goes on to say that membership through any of several facts belongs in a segment's `measures` list; see [Cohorts and segments](#cohorts-and-segments). ## Lists and the quoting gotcha `in_list` and `exclude_list` take one string. Items are split on commas, trimmed and lower-cased. A comma inside single or double quotes does not split, and the quotes stay part of the item. ```json {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""} ``` renders, on every page of this cookbook, as ```sql LOWER(T1."i_category") IN ('men', 'children', '"books,com"') ``` So `"books,com"` matches a category whose stored value includes the quotes. Quote an item only when the warehouse value really contains a comma and the quotes, otherwise list it bare. The shorthand strips quotes from list items before the planner sees them, so `category in (Books, "Music, Live")` becomes three values; use the JSON spec for list items that contain commas. ## Dates: fixed and relative Fixed formats include `2024-01-31`, `2024/01/31`, `31-01-2024`, `Jan 31 2024`, `31 Jan 2024` and the same with a time (`2024-01-31 09:30:00`, `2024-01-31T09:30:00`). Relative values are one to three digits and a unit: `28d`, `3m`, `1y`, also `h`, `w`, `q`. They mean "that many units ago"; `w`, `m`, `q` and `y` snap to the start of the unit, or to its end when used as `value_end`. ```json "filters": [ {"field": "date", "predicate": "greater_than_or_equal_to", "value": "28d"} ] ``` ```json "filters": [ {"field": "date", "predicate": "between", "value": "3m", "value_end": "1m"} ] ``` The second reads: from the first day of the month three months ago to the last day of last month. `between` needs both `value` and `value_end`. An unparseable date is a `Planner::ResolutionError` `cannot parse date ...`. ## String predicates `contains`, `starts_with`, `ends_with` and their `does_not_` forms render as `LIKE` on a lower-cased column. ```json {"spec": {"projections": [{"field": "ws_net_paid"}], "filters": [{"field": "category", "predicate": "starts_with", "value": "super"}]}} ``` ```sh zsql sql --expr "web net paid, category starts with super" ``` ```sql SELECT sum(T0."ws_net_paid") AS "Web Net Paid" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") LIKE 'super%' ``` `contains` gives `LIKE '%super%'`, `does_not_contain` gives `NOT LIKE`, `keyword` with `a, b` gives `(col LIKE '%a%' OR col LIKE '%b%')`. `equals` on a string is `LOWER(col) = 'books'`. ## A measure filter becomes HAVING There is no separate syntax. `{"field": "ws_net_paid", "predicate": "greater_than", "value": "100"}` is a measure filter because `ws_net_paid` is a measure, and the planner places it in the `HAVING` of the node that groups: ```sql HAVING sum(T0."ws_net_paid") > 100 ``` The full statement, with the `HAVING` inside an aggregation CTE next to the `WHERE` for the category list, is on [Period over period](#period-over-period). Measures accept `equals`, `does_not_equal`, `between` and the four comparisons; `top_n` on a measure is refused with `Predicate top n can only be applied to a dimension field`. ## Top-n Web revenue for the five best categories by web revenue, among three candidates. ```json {"spec": { "projections": [{"field": "ws_net_paid"}, {"field": "category"}], "filters": [ {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}, {"field": "category", "predicate": "top_n", "value": "5", "top_n_measure": "ws_net_paid"} ] }} ``` ```sh zsql sql --expr "web net paid, category, category in (men, children), category top 5 by web net paid" ``` ```sql WITH ag5c9bb8d550197b3a219ae50f0b427905 AS ( SELECT T1."i_category" AS "dim30d09b7", sum(T0."ws_net_paid") AS "msr0c0d023" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ORDER BY sum(T0."ws_net_paid") desc LIMIT 5 ) SELECT sum(T0."ws_net_paid") AS "Web Net Paid", T1."i_category" AS "Category" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk INNER JOIN ag5c9bb8d550197b3a219ae50f0b427905 AS A2 ON T1."i_category" = A2.dim30d09b7 WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ``` - **A ranking CTE picks the members.** `ag5c9b…` aggregates the measure by the dimension, applies the same `WHERE` as the query, orders by the measure descending and takes `LIMIT 5`. This is the only `LIMIT` the planner emits; the spec's `limit` key is carried, not rendered. - **The query is INNER JOINed to it on the dimension.** The main query then aggregates normally, so the projected measure is computed over the full rows of the winning categories. - **The measure is optional.** Omit `top_n_measure` to rank by row count. The measure can also be packed into the value as `"5:ws_net_paid"`. ## Variations - **Bottom five.** There is no `bottom_n`; project the measure with `"order_by": "asc"` and apply your own cut, or filter with `less_than` against a threshold. - **Null handling.** `is_null` and `is_not_null` take no value: `{"field": "date", "predicate": "is_not_null"}`, shorthand `date is not null`. - **Exclude a list.** `{"field": "category", "predicate": "exclude_list", "value": "books, music"}`, shorthand `category not in (books, music)`, renders as `NOT IN`. ## Next steps - [Filters](#filters) for the predicate-by-type table and validation messages - [Cohorts and segments](#cohorts-and-segments) for OR-across-facts membership - [Shorthand](#shorthand-expressions) ### Level of detail Some measures must be computed at a grain other than the one the query groups by: a grand total repeated on every row, a total that ignores the category filter the caller applied, or a median of per-ticket sums. 0sql models these as exclusion and inclusion measures in YAML. At request time they are plain fields; the planner notices the rule and adds the extra aggregation. This page shows the model shapes, a spec that uses them, the shape of the statement, and when the request-time `contribute` decorator is the better tool. Model reference: [Exclusions](#exclusions) and [Inclusions](#inclusions). ## Exclusion measures in the model An exclusion rule removes dimensions from a measure's grouping and says what to do with filters on them. The all-categories denominator from the [Exclusions](#exclusions) page, written for TPC-DS web sales: ```yaml - type: measure name: Web Net Paid data_type: decimal expression: sql: sum(ws_net_paid) - type: measure name: Web Net Paid (All Categories) data_type: decimal exclusion_type: exclude exclusions: - type: dimension filter: ignore # a filter on Category does not apply to this measure entities: - Category expression: sql: sum(ws_net_paid) ``` Two independent knobs: `exclusion_type` (`exclude`, `exclude_all_except`, `exclude_all`) controls which dimensions may group the measure; `filter` (`apply`, `ignore`, `only`) controls what filters on those dimensions do. The running TPC-DS project carries this one, which drops every dimension of the `Date` table and ignores date filters: ```yaml - type: measure name: Store Level Total Sales description: Exclude date_dim dimensions and ignore filter data_type: decimal exclusion_type: exclude exclusions: - type: table filter: ignore entities: [Date] expression: sql: sum(ss_net_paid) ``` ## Inclusion measures in the model An inclusion rule adds dimensions for an inner aggregation and rolls the result up with a second aggregate. From the running project: ```yaml - type: measure name: Median Store Order Size description: This is the median order total per sale data_type: decimal inclusions: filter: apply aggregation: percentile_cont(0.5) WITHIN GROUP (ORDER BY @exp) dimensions: [Store Ticket number] expression: sql: sum(ss_net_paid) ``` Inner: `sum(ss_net_paid)` grouped by the query's dimensions plus `Store Ticket number`. Outer: the median of those per-ticket sums, grouped by the query's dimensions only. `@exp` stands for the inner result. ## The request The exclusion measure beside the plain measure, plus a calculation that divides them. ```json {"spec": { "projections": [ {"field": "category", "alias": "Category"}, {"field": "Web Net Paid", "alias": "Web Net Paid"}, {"field": "Web Net Paid (All Categories)"}, {"alias": "Share", "calculation": true, "data_type": "decimal", "sql": "[Web Net Paid]@m / nullif([Web Net Paid (All Categories)]@m, 0)"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men, children"}] }} ``` ```sh zsql sql --expr "category, [Web Net Paid], [Web Net Paid (All Categories)], Share = [Web Net Paid]@m / nullif([Web Net Paid (All Categories)]@m, 0), category in (men, children)" ``` Nothing in the spec says "exclusion". The field name is enough; the rule travels with the measure. Field names with parentheses are fine in brackets. ## The shape of the SQL We have no captured statement to quote for this exact request, so here is the shape rather than the text. The planner runs an `exclusion` phase (you can see it in `POST …/explain` under `phases`) that forks the exclusion measure into its own aggregation: one CTE sums `ws_net_paid` by `i_category` with the category `WHERE`, a second CTE sums `ws_net_paid` with no `GROUP BY` on category and, because `filter: ignore`, no category `WHERE`. The final `SELECT` joins the coarser aggregation back to the per-category rows (a single total row joins every row, as the `CROSS JOIN` does on the contribute page), projects both columns, and computes `Share` from the two aggregated columns at the merge, exactly where the `Bucket` calculation sits on [Complex measures](#complex-measures). Every category row carries the same denominator, so `Share` is each category's fraction of the all-category total even though the query is filtered to two categories. For the inclusion measure the shape is nested instead of parallel: an inner aggregation grouped by the query's dimensions plus `Store Ticket number`, wrapped by an outer `SELECT` that applies `percentile_cont(0.5) WITHIN GROUP (ORDER BY …)` grouped by the query's dimensions. Request it with `{"field": "Median Store Order Size"}` next to `Item Category` and it joins the result like any other measure column. ## Exclusion measure or contribute decorator Both give a percent of total. They differ in who decides and in what the denominator respects. | | exclusion measure in YAML | `contribute` decorator in the spec | |---|---|---| | defined by | the modeler, once | the caller, per request | | denominator | whatever the rule says: a table, a universe, a named dimension; filters applied, ignored or exclusively applied | the total of the measure, partitioned by `partition_refs`; filters on the partition dimensions ignored by default (`ignore_partition_filters`) | | discoverable | yes, it is a field (`GET …/fields`) | no, it is a decorator | | composable | in other compound measures and in calculations by `[Name]@m` | only as the projection it decorates | | security | policies trigger on it like any field; its own rule on the context dimension outranks a resolved filter | the decorated measure's field triggers policies | | when | the ratio is a governed number (share of wallet, percent of plan, index to total) | an ad hoc "what fraction is this" on any measure | Rule of thumb: if the denominator needs an opinion about filters, tables or universes, model it. If the caller just wants each row divided by the sum of the rows it can see, decorate it. See [Share and windows](#share-and-windows) for the decorator's SQL. ## Variations - **Grand total on every row.** `exclusion_type: exclude_all` with no `exclusions` list gives a measure computed as one value regardless of the query's dimensions; a query for `category, Web Net Paid, Grand Total` repeats the total on each row. - **Respect the filter, drop the grouping.** `filter: apply` instead of `ignore` on `Web Net Paid (All Categories)` keeps the category `WHERE` in the coarser CTE, so the denominator is the total of the two filtered categories and the shares sum to 100%. - **Median by month.** Project `{"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}` with `Median Store Order Size`; the inner aggregation groups by month and ticket, the outer by month. ## Next steps - [Exclusions](#exclusions) and [Inclusions](#inclusions) - [Share and windows](#share-and-windows) - [Complex measures](#complex-measures) ### Period over period You want each month's web revenue next to the previous month's, or the growth between them, without writing a self-join. The `temporalize` decorator on a measure asks the planner for the same measure shifted by one period. It needs a date projection truncated to the matching grain in the same request, because the shift is expressed on that date. Reference: [Projections and decorators](#projections-and-decorators). ## Transforms | JSON `transform` | shorthand | pair with | |---|---|---| | `year_over_year` | `yoy(x)`, `yoy_pct(x)` | `year(date)` | | `quarter_over_quarter` | `qoq(x)`, `qoq_pct(x)` | `quarter(date)` | | `month_over_month` | `mom(x)`, `mom_pct(x)` | `month(date)` | | `week_over_week` | `wow(x)`, `wow_pct(x)` | `week(date)` | | `day_over_day` | `dod(x)`, `dod_pct(x)` | `day(date)` | `"percent_change": true` (the `_pct` forms) returns the relative change instead of the prior-period value. The derived alias is `LM(Web Net Paid)`, `LY(…)`, `LQ(…)`, `LW(…)` or `D/D(…)`, prefixed with `%` when `percent_change` is set. Set your own `alias` to override it. ## Month over month with a measure filter Previous month's web revenue per month, for three categories, keeping only months above 100. ```json {"spec": { "projections": [ {"field": "date", "order_by": "asc", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "ws_net_paid", "decorators": [{"type": "temporalize", "transform": "month_over_month"}]} ], "filters": [ {"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}, {"field": "ws_net_paid", "predicate": "greater_than", "value": "100"} ] }} ``` ```sh zsql sql --expr "month(date) asc, mom(web net paid), category in (men, children), web net paid > 100" ``` ```sql WITH ag47a9128558e109f9cf0b93c01992e0e7 AS ( SELECT (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) AS "dim5b35e57", sum(T0."ws_net_paid") AS "__msrlm_51898eff8" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY (DATE_TRUNC('month', T1."d_date")::DATE + '1 month'::interval) HAVING sum(T0."ws_net_paid") > 100 ) SELECT A0.dim5b35e57 AS "Month(Date)", A0.__msrlm_51898eff8 AS "LM(Web Net Paid)" FROM ag47a9128558e109f9cf0b93c01992e0e7 A0 ORDER BY A0.dim5b35e57 asc ``` ## Why the SQL looks like this - **A shifted-date aggregation CTE.** `ag47a9…` groups web revenue by `DATE_TRUNC('month', d_date) + '1 month'`. The sum for January is stored under February's key, so when the final `SELECT` reads the row for February it gets January's value: that is `LM(Web Net Paid)`. - **The date projection supplies the grain.** The `truncate` on `date` is what the shift is applied to. Without a month-truncated date projection there is nothing to shift by. - **HAVING stays inside the CTE.** The measure filter `ws_net_paid > 100` is a filter on an aggregate, so it is `HAVING sum(…) > 100` in the node that groups, which is the shifted CTE. The dimension filter on category is a `WHERE` in the same node. See [Filters and top-n](#filters-and-top-n) for the placement rule. - **Display-only transform.** Only the shifted measure is projected, so the final `SELECT` reads the CTE directly and orders by the month. Project the plain `ws_net_paid` too and the planner also needs the unshifted aggregation; the two are joined on the month key so current and prior sit on one row. ## Variations - **Year over year, as a percentage.** `zsql sql --expr "yoy_pct(net paid), year(date)"`, or in JSON `{"field": "net_paid", "decorators": [{"type": "temporalize", "transform": "year_over_year", "percent_change": true}]}` with `{"field": "date", "decorators": [{"type": "truncate", "grain": "year"}]}`. The shifted CTE adds one year instead of one month, and the final `SELECT` computes the change relative to the prior value, and the derived alias gets a `%` prefix. - **Prior and current together.** Projections `month(date)`, `web net paid`, `mom(web net paid)` give three columns per month. In the shorthand: `zsql sql --expr "month(date) asc, web net paid, mom(web net paid)"`. - **Keep the filter off the comparison.** Move `web net paid > 100` out of `filters` and apply the threshold in your own code if you want the prior-month column to include months under 100. The `HAVING` applies to the shifted node, so it drops prior months by their own value, not by the current month's. ## Next steps - [Projections and decorators](#projections-and-decorators) for every decorator attribute - [Share and windows](#share-and-windows) for running totals over the same month axis - [Shorthand](#shorthand-expressions) ### Security context Row-level security in 0sql is part of planning. A policy in `security.yml` names the fields that trigger it, the dimension that holds the allowed values, and where in the caller's context those values come from. Your application builds a `context` object for the user making the request and sends it beside the spec. The planner resolves the allowed values from it and adds a `WHERE` (filter rows) or a `CASE` (mask values) to the statement. Nothing from the context is stored. Reference: [Security context](#security-context) and [Security](#row-level-security). ## The policy A call-center project tags `Country` and `Employees` on the `Call Center` table with `pii`, and restricts them to the call centers a user's groups are tagged with. ```yaml # models/common/tbl.call_center.yml (fragment) - type: dimension name: Call Center ID data_type: string expression: sql: cc_call_center_id - type: dimension name: Country data_type: string tags: [pii] expression: sql: cc_country - type: dimension name: Employees data_type: integer tags: [pii] expression: sql: cc_employees ``` ```yaml # security.yml policies: - name: Call center rows mode: filter_data triggers: field_tags: [pii] context_dimension: Call Center ID permission_resolution: source: groups value_from: tag tag_key: call_center_id unresolved: allow bypass: system_admin: true project_admin: false ``` Read it as: when a query projects a field tagged `pii`, collect the values of `call_center_id:` tags across the user's groups, and keep only rows whose `Call Center ID` is one of them. A system admin skips the policy. If the context yields no values the policy is skipped (`unresolved: allow`); with the default `deny` the query would return no rows (`WHERE 1 = 0`). ## The context your app builds Per request, from your own session or identity provider: ```json { "email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [ {"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]} ] } ``` Every key is optional; unknown keys are rejected. `tags` are the user's own `key:value` tags for `source: user` policies; `groups` carry names and tags for `source: groups` policies. The `cat:men` tag is ignored by this policy because its `tag_key` is `call_center_id`. ## filter_data ```sh curl -s https://app.0sql.io/projects/tpcds/branches/main/sql \ -H "Authorization: Bearer zqk_..." \ -H "Content-Type: application/json" \ -d '{"spec": {"name": "security_test", "projections": [{"field": "country", "alias": "Country"}]}, "context": {"email": "tank@matrix.com", "system_admin": false, "project_admin": false, "tags": [], "groups": [{"name": "CC-TMNT", "tags": ["call_center_id:TMNT"]}, {"name": "CC-AVCL", "tags": ["call_center_id:AVCL", "cat:men"]}]}}' ``` With `zsql`: `zsql sql --expr "country as Country" --context ctx.json`. ```sql SELECT T0."cc_country" AS "Country" FROM call_center T0 WHERE LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl') ``` - **The policy fired** because `Country` carries the `pii` tag. - **Allowed values came from group tags.** `TMNT` and `AVCL` are the values of the `call_center_id:` tags; the comparison is lower-cased on both sides. - **The context dimension need not be projected.** `Call Center ID` is reached through the model; if it cannot be reached from the query's tables the request fails with `Planner::SecurityPolicyError` rather than returning unfiltered rows. ## mask_data Change the policy to `mode: mask_data` (with `mask_value: "#######"`) and project `employees`, an integer, with the same context: ```sql SELECT CASE WHEN LOWER(T0."cc_call_center_id") IN ('tmnt', 'avcl') THEN T0."cc_employees" ELSE NULL END AS "Employees" FROM call_center T0 ``` Rows stay. The triggered field is replaced where the context dimension is outside the allowed set. The mask is `mask_value` for string fields and `NULL` for numeric fields, so an integer column is never handed a string literal. ## system_admin bypass Send the same filter_data request with `"system_admin": true` in the context and the statement is the plain projection: ```sql SELECT T0."cc_country" AS "Country" FROM call_center T0 ``` `bypass.system_admin` defaults to `true`; `bypass.project_admin` defaults to `false`. Set both to `false` on a policy that should apply to everyone. ## Omitting the context On a branch that has any policy, a request without `context` is refused before planning: ```json {"error": {"class": "ContextRequired", "message": "this branch has security policies; a context is required to plan"}} ``` HTTP 400. On a branch with no policies the context is optional and the security phase is skipped. ## Variations - **Two policies, two predicates.** A second policy keyed on `tag_key: cat` and triggered by another tag adds `AND LOWER(…) IN ('men')` when both tagged fields are projected; the `cat:men` tag in the context above is what unlocks it. - **Self-service by email.** `source: user`, `value_from: email` resolves the context's `email` as the single allowed value: `WHERE LOWER(employee_email) IN ('tank@matrix.com')`. The user's own `tags` work the same way with `value_from: tag`. - **Group names as values.** `source: groups`, `value_from: name` uses the group names (`CC-TMNT`, `CC-AVCL`) directly; no `tag_key`. ## Next steps - [Security context](#security-context) for the resolution rules - [Security](#row-level-security) for every policy key - [Level of detail](#level-of-detail): a measure's exclusion rule on the context dimension outranks a resolved filter ### Share and windows Two families of measure decorators reshape an aggregated number against its neighbours. `contribute` divides each row's measure by a total computed in a separate CTE, which gives percent of total. `window` wraps the aggregate in a SQL window function (`sum`, `avg`, `lag`, `rank`, …) over the query's own rows, for running totals, moving averages and offsets. Both apply to measures only. Reference: [Projections and decorators](#projections-and-decorators). ## Percent of total Each category's share of web revenue. ```json {"spec": { "projections": [ {"field": "ws_net_paid", "decorators": [{"type": "contribute"}]}, {"field": "category"} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] }} ``` ```sh zsql sql --expr "pct(web net paid), category, category in (men, children)" ``` ```sql WITH ag61495e144c0531efb642c4545df2e6c5 AS ( SELECT sum(T0."ws_net_paid") AS "msr9f2f6d5", T1."i_category" AS "dim841154e" FROM web_sales T0 JOIN item T1 ON T0.ws_item_sk = T1.i_item_sk WHERE LOWER(T1."i_category") IN ('men', 'children', '"books,com"') GROUP BY T1."i_category" ), agbe9574d2ae0d53a3a40c210cf88ff8c4 AS ( SELECT sum(T0."ws_net_paid") AS "__msrtotal_51898eff8" FROM web_sales T0 ) SELECT A0.msr9f2f6d5 / (A1.__msrtotal_51898eff8*1.00) AS "% Web Net Paid of Total", A0.dim841154e AS "Category" FROM ag61495e144c0531efb642c4545df2e6c5 A0 CROSS JOIN agbe9574d2ae0d53a3a40c210cf88ff8c4 A1 ``` - **The per-row aggregation is the first CTE.** `ag6149…` is web revenue by category, with the category filter. - **The total is a second CTE with no GROUP BY.** `agbe95…` sums the whole fact. It carries no `WHERE`: `ignore_partition_filters` defaults to `true`, so filters on the dimensions that define the share (here `Category`, the only projected dimension) are left out of the denominator. The share is relative to all categories, and the three visible rows do not sum to 100%. Set `"ignore_partition_filters": false` on the decorator to make them. - **CROSS JOIN, then divide.** One total row is cross-joined to every category row and the final `SELECT` computes `msr / (total*1.00)`. The `*1.00` forces decimal division. - **The alias is derived.** `% Web Net Paid of Total`, unless you set `alias`. ## Share within a category Project `category` and `product_name`, and tell `contribute` which projected dimensions partition the total. ```json {"spec": { "projections": [ {"field": "category"}, {"field": "product_name"}, {"field": "ws_net_paid", "alias": "Share of Category", "decorators": [{"type": "contribute", "partition_refs": ["category"]}]} ] }} ``` ```sh zsql sql --expr "category, product name, pct(web net paid, category) as Share of Category" ``` No captured statement to quote for this one. The shape follows from the previous statement: the total CTE now groups by `i_category`, and the final `SELECT` joins it on the category instead of cross-joining a single row, so each product is divided by its own category's total. `partition_refs` must name dimensions that are projected; a ref that is not projected is dropped. ## Running sum and moving average ```json {"spec": { "projections": [ {"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "ws_net_paid", "decorators": [{"type": "window", "mode": "running", "function": "sum"}]}, {"field": "ws_net_paid", "decorators": [{"type": "window", "mode": "moving", "function": "avg", "size": 7}]} ] }} ``` ```sh zsql sql --expr "month(date), running_sum(web net paid), moving_avg(web net paid, 7)" ``` The ordering defaults to `{"-1": "asc"}`, meaning the first projected date, ascending; `partition_refs` is empty, so the window runs over the whole result. The window function wraps the aggregate, as in the lag statement below: a moving average of size 14 renders in the engine's tests as `avg(sum(x)) OVER ( ORDER BY ROWS BETWEEN 14 PRECEDING AND CURRENT ROW)`, and a running sum is the same pattern with `sum(sum(x))` and no `ROWS` frame. There is no separate CTE because the window reads the same grouped rows. Aliases: `Running Sum(Web Net Paid)` and `Moving Avg(Web Net Paid)`. ## Lag with an explicit partition and order Two categories back within each month, categories ordered descending. ```json {"spec": { "projections": [ {"field": "date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "category"}, {"field": "ws_net_paid", "decorators": [{"type": "window", "mode": "running", "function": "lag", "offset": 2, "partition_refs": ["date"], "order_by": {"category": "desc"}}]} ], "filters": [{"field": "category", "predicate": "in_list", "value": "men ,children, \"books,com\""}] }} ``` ```sql SELECT DATE_TRUNC('month', T1."d_date")::DATE AS "Month(Date)", T2."i_category" AS "Category", lag(sum(T0."ws_net_paid"), 2) OVER (PARTITION BY DATE_TRUNC('month', T1."d_date")::DATE ORDER BY T2."i_category" desc) AS "Running Lag(Web Net Paid)" FROM web_sales T0 JOIN date_dim T1 ON T0.ws_sold_date_sk = T1.d_date_sk JOIN item T2 ON T0.ws_item_sk = T2.i_item_sk WHERE LOWER(T2."i_category") IN ('men', 'children', '"books,com"') GROUP BY DATE_TRUNC('month', T1."d_date")::DATE, T2."i_category" ``` `partition_refs` and `order_by` are maps over projected dimension uids: `date` partitions, `category desc` orders. The `offset` becomes the second argument of `lag`. The shorthand `running_lag(web net paid, date)` sets the partition but has no syntax for `offset` or `order_by`; use JSON when you need them. ## Variations - **Rank within a partition.** `{"type": "window", "mode": "running", "function": "rank", "partition_refs": ["date"], "order_by": {"category": "desc"}}`; `dense_rank`, `row_number`, `percent_rank`, `cume_dist` and `ntile` (with `"buckets": 4`) work the same way. `order_by` keys are projected dimension uids; a key that is not a projected dimension is dropped. - **Running sum per category.** `zsql sql --expr "month(date), category, running_sum(web net paid, category)"` adds `partition_refs: ["category"]`, so each category's total restarts. - **Share of a total from the model.** When the denominator must survive filters your callers add, define it as an exclusion measure in YAML instead of using `contribute`; see [Level of detail](#level-of-detail). ## Next steps - [Projections and decorators](#projections-and-decorators) for the window attributes and their validation - [Period over period](#period-over-period) - [Level of detail](#level-of-detail) ### TPC-DS Tutorial Complete walkthrough of the TPC-DS benchmark semantic model. ## Overview The TPC-DS Tutorial is a nearly complete model of the TPC-DS benchmark system. It demonstrates best practices for modeling a complex data warehouse with 0sql. ## Project Structure ``` tpcds-tutorial/ ├── project.yml # Project configuration ├── datasources.yml # Datasource definitions ├── models/ # Semantic model files │ ├── catalog/ # Catalog channel models │ ├── common/ # Shared dimension tables │ ├── inventory/ # Inventory models │ └── store/ # Store channel models └── tests/ # Planner tests, run on every deploy ``` ## Key Components ### Datasource Configuration ```yaml tpcds: name: TPC-DS (Postgres) adapter: postgres tier: warm host: localhost port: 5432 database: tpcds schema: public username: tpcds ``` `adapter` and `tier` are what the planner uses, and `name` is what the response calls the datasource. The connection keys are carried as metadata for your own application; 0sql never uses them, and `zsql deploy` strips any secret before building the archive. ### Table Models #### Fact Tables **Store Sales** (`models/store/tbl.store_sales.yml`): ```yaml name: Store Sales physical_name: store_sales datasource: tpcds cost: 100 # Fact table: higher cost fields: - type: dimension name: Store Ticket number data_type: integer expression: primary_key: true sql: ss_ticket_number - type: measure name: Store Quantity data_type: integer expression: sql: sum(ss_quantity) - type: measure name: Store Sales Price data_type: decimal expression: sql: sum(ss_sales_price) ``` #### Dimension Tables **Store** (`models/store/tbl.store.yml`): ```yaml name: Store physical_name: store datasource: tpcds cost: 10 # Dimension table: lower cost fields: - type: dimension name: Store Name data_type: string expression: sql: s_store_name ``` ### Relationships **Store Relationships** (`models/store/rel.store.yml`): ```yaml datasource: tpcds # Store Sales to Date Dimension store_sales_sold_date: left: Store Sales right: Date sql: left.ss_sold_date_sk = right.d_date_sk cardinality: many_to_one # Store Sales to Customer Dimension store_sales_customer: left: Store Sales right: Customer sql: left.ss_customer_sk = right.c_customer_sk cardinality: many_to_one # Store Sales to Store Dimension store_sales_store: left: Store Sales right: Store sql: left.ss_store_sk = right.s_store_sk cardinality: many_to_one ``` ## Key Patterns ### 1. Cost Configuration - **Dimension tables**: `cost: 10` (lower cost = preferred) - **Fact tables**: `cost: 100` (higher cost = used when needed) ### 2. Primary Keys Mark unique identifiers as primary keys: ```yaml - type: dimension name: Store Ticket number expression: primary_key: true sql: ss_ticket_number ``` ### 3. Measure Aggregation All measures include aggregation: ```yaml - type: measure name: Store Quantity expression: sql: sum(ss_quantity) ``` ### 4. Relationship Cardinality Most relationships are `many_to_one`: ```yaml store_sales_store: left: Store Sales right: Store sql: left.ss_store_sk = right.s_store_sk cardinality: many_to_one ``` ## Testing Test files validate SQL generation: ```yaml name: Store Sales Positive Values projections: - Store Sales Price - Store Quantity assert_sql: "SELECT sum(T0.\"ss_quantity\") AS \"Store Quantity\", sum(T0.\"ss_sales_price\") AS \"Store Sales Price\" FROM store_sales T0" ``` Tests run on every `zsql deploy` and can be re-run with `zsql test`. A failing test exits non-zero, which is what you want in CI. ## Lessons Learned 1. **Organize by domain** - Group related tables (catalog, store, common) 2. **Use appropriate costs** - Lower for dimensions, higher for facts 3. **Mark primary keys** - Helps with query optimization 4. **Define relationships** - Enable cross-table queries 5. **Test the planner** - Assert the SQL a projection produces in `tests/*.yml` ## Next Steps - [Explore the project](https://github.com/stratasite/tpcds-tutorial) - [Learn star schema pattern](#star-schema-pattern) - [Try customer 360 recipe](#customer-360-recipe) --- ## Troubleshooting What the common errors mean and how to fix them, from a deploy the loader rejects to a query the planner refuses. 0sql refuses a query it cannot answer correctly instead of returning a wrong number; the message names the field or combination it could not resolve, and that is the thing to fix in the model, not something to work around in the request. Every HTTP error has the same shape, and `zsql` prints it as `: `: ```json {"error": {"class": "Planner::ResolutionError", "message": "No universe can resolve the query within datasource tpcds."}} ``` ## Quick reference | Symptom | Status and class | Start here | |---|---|---| | `no project.yml here or above` | local `zsql` | [Deploy errors](#deploy-errors) | | `Error in models/...: ...` | 400 `DeployError` | [Model validation](#model-validation) | | `the archive holds no project.yml` | 400 `DeployError` | [Deploy errors](#deploy-errors) | | `tests failed` | exit code only; deploy was 200 | [Deploy errors](#deploy-errors) | | `this branch has security policies; a context is required to plan` | 400 `ContextRequired` | [Plan errors](#plan-errors) | | `No field named 'x' in this model.` | 422 `Semantic::NotFound` | [Fuzzy field corrections](#fuzzy-field-corrections) | | `No universe can resolve the query within datasource ...` | 422 `Planner::ResolutionError` | [Query resolution](#query-resolution) | | `Security policy context dimension '...' is not reachable` | 422 `Planner::SecurityPolicyError` | [Query resolution](#query-resolution) | | `an API key is required` / `this API key is not valid` | 401 `Unauthorized` | [Auth errors](#auth-errors) | | `query key ... is not granted ...` | 403 `Forbidden` | [Auth errors](#auth-errors) | | A field you did not ask for appears in the SQL | `corrections` in the response | [Fuzzy field corrections](#fuzzy-field-corrections) | | Numbers too high | model, not request | [Numbers look wrong](#numbers-look-wrong) | ## Deploy errors `zsql deploy` tars the project directory and POSTs it; `zsql check` does the same against `/validate` and stores nothing. Both fail with a 400 `DeployError` whose message says which stage refused. ### `no project.yml here or above; run zsql init to start a project` Local. `zsql` walks up from the current directory looking for `project.yml`. Run it inside the project, or pass `--project `. ### `the body must be a tar.gz of the project directory`, `unpacking the archive: ...`, `the archive holds no project.yml` Only reachable when you POST to `/deploy` yourself. The body must be a gzip tar with `project.yml` at its root or inside one top-level folder. `zsql deploy` builds this for you. ### `Error in : ` The loader rejected a file. The message is one of the [model validation](#model-validation) messages below; fix the file and run `zsql check` until it is clean, then `zsql deploy`. Nothing is deployed when any file fails: the branch keeps its previous model. ### `snapshot has N integrity problems:` A compiled model that references something missing. Each line is of the form `expression references missing field`, `join left table is missing`, `field mapped twice on table `, `table snapshot date is not a dimension`, `partition

dimension is missing or not a dimension`, or `duplicate uid `. Almost always a rename: a table or field was renamed in one file and still referenced by its old name in a relation, a `snapshot:`, a partition or a formula. ### `query keys are read only; deploy with your personal key` 403. The key in `.zsql` or `ZSQL_API_KEY` starts with `zqk_`. Deploy with a `zsk_` key. See [Auth errors](#auth-errors). ### Deploy succeeded but `zsql deploy` exited with `tests failed` The server returns 200 even when tests fail; the model is live. `zsql` prints `PASSED`, `FAILED` or `ERROR` per test with `--- generated` and `--- expected` blocks, then exits non-zero. A failing `assert_sql` usually means the model changed and the expected SQL is stale: compare, update the test, redeploy. `ERROR` means the test's projections no longer plan (a renamed field) or its `assert_regex` does not compile. See [Tests](#tests). ### Warnings A deploy can succeed with `warning:` lines. They are worth fixing: | Warning | Meaning | |---|---| | `: table has no cost; 0 assumed` | Add `cost:` so routing prefers the right table. | | `: table\|field\|join : unknown key '' is ignored` | A typo in a YAML key. The setting is not applied. | | `: partition on : predicate

is not supported` / `needs a number or date dimension` | The partition will not rank the table. See [Partitions](#partitions). | | `formation: root reaches by two routes of cost : ...` | Two equal-cost join paths; one was chosen. Set costs so the choice is deliberate. | | `formation: ambiguous: from the field "" is reachable at cost through A and B; A is used` | Same field reachable through two dimension tables. Set costs or split the dimension. | | `policy context dimension not found; policy is inert` | The policy's `context_dimension` was removed or renamed. The policy no longer protects anything. | | `project.yml not found; the directory name stands in for the project` | Add a `project.yml` with a stable `uid`. | ## Model validation Loader messages arrive as `Error in : ` followed by one of these. The ones people hit most: | Message | Fix | |---|---| | `datasources.yml is required` / `datasources.yml defines no datasource` | Add at least one datasource keyed by a short id. | | `Datasource errors: adapter '' is not supported` | Use one of the supported adapters. See [Datasources](#datasources). | | `Datasource errors: tier '' should be hot, warm or cold ()` | Fix `tier:`. | | `datasource key is required in table file.` | Every `tbl.*.yml` and `rel.*.yml` names its `datasource:`. | | `datasource with uid or name not found in branch.` | The `datasource:` value does not match a key in `datasources.yml`. | | `Table errors: Name has already been taken ()` | Two table files share a `name`. | | `Table errors: Physical name can't be blank ()` | Add `physical_name:`. | | `fields must be a list ()` / `each field of must be a mapping` | YAML shape. Indentation (spaces, not tabs), a missing `-` or colon. | | `Field errors: type should either be a dimension or measure` | `type:` is required on every field. | | `Field errors: Data type '' is not included in the list` | See [Data Types](#data-types). | | `Field '' is missing an expression node.` / `Expression errors for : Sql can't be blank` | Every field needs `expression: { sql: ... }`. | | `Expression errors for : Sql measure should have an aggregation function` | A measure's `sql` must aggregate: `sum(ss_net_paid)`, not `ss_net_paid`. | | `Expression errors for : Column refs dimensions should reference a table column` | A dimension's expression must read a column of its table. | | `Expression errors for : Field is already mapped to this table.` | The same field twice in one table. The same name on another table is not an error: it is one concept. | | `Field errors: Grains contains invalid values: ` / `Snapshot '' is not included in the list` | The message lists the accepted values. See [Snapshot measures](#snapshot-measures). | | `Field errors: Exclusion rule Type should be one of: dimension, table, universe` | See [Exclusions](#exclusions). | | `JoinDef errors for : Cardinality can't be blank` / `Cardinality '' is not included in the list` | `many_to_one`, `one_to_many` or `one_to_one`. Many-to-many is not supported; model a junction table with two relations. See [Cardinality](#cardinality). | | `JoinDef errors for : Sql join sql should be of the format left.column = right.column` | Only equi-joins, written with the `left.` and `right.` prefixes. | | `Left table '' not found in datasource ` / `Right table ...` | Relations never cross datasources, and names must match the table files' `name`. | | `Join: <->, already exists.` | Two relations between the same pair of tables. | | `Circular import detected: a.yml -> b.yml -> a.yml` / `Import not found: ` | See [Imports](#imports). | | `could not find snapshot dimension in branch.` | The table's `snapshot:` must name a date dimension. | | `Partition errors: Dimension must exist ()` / `Predicate is not included in the list ()` | See [Partitions](#partitions). | | `[]: could not find measure named: ` / `could not find dimension named: ` | A `[Name]@m` or `[Name]@d` reference in a formula does not exist. Names are matched case-insensitively; check spelling. | | `[] on : Sql Referenced dimensions in the formula could not be located in a single Universe.` | A compound measure references a dimension not every component can reach. Move the conditional part into a standard measure on the fact that has the dimension. See [Compound Measures](#compound-measures). | | `[] on
: Sql Could not find a datasource which had all of the required measures. Cross datasource queries are not supported.` | Components of a compound measure live in different datasources. Define the missing measure in one of them under the same name. | | `[]: Inclusions Dimensions [...] could not be found in a universe with this measure.` | See [Inclusions](#inclusions). | | `Policy '': context_dimension '' not found in this branch.` / `trigger field '' not found in this branch.` | Names in `security.yml` must match a field in the branch. See [Security](#row-level-security). | | `Unknown mode ''. Use mask_data or filter_data.` / `Policy '': unresolved '' should be deny or allow` | Fix the value. | | `Test must have a name` / `Test must have at least one field in 'projections'` / `Test must have at least one assertion (assert_sql or assert_regex)` | See [Tests](#tests). | ## Plan errors `POST .../sql`, `/explain` and `/explore` answer 400 for a malformed request, 404 when the branch is not deployed, and 422 when the request is well formed but the model cannot answer it. | Status | Class | Message | What to do | |---|---|---|---| | 400 | `Invalid` | `give a spec or an expr` | The body needs `spec` or `expr`. | | 400 | `Shorthand` | the parse error | The `expr` line did not parse. See [Shorthand](#shorthand-expressions). | | 400 | `ContextRequired` | `this branch has security policies; a context is required to plan` | Send a `context`. See [Security](#row-level-security). | | 404 | `NotFound` | `no deployment for project branch ` | Deploy that branch, or check `--branch` (the default is the checked-out git branch). `zsql list` shows what is deployed. | | 422 | `Semantic::NotFound` | `No field named '' in this model.` optionally followed by ` Did you mean: A, B, C?` | No field matches by name, uid or synonym. `zsql fields ` searches. See [Fuzzy field corrections](#fuzzy-field-corrections). | | 422 | `Semantic::Ambiguous` | `Field '' is ambiguous ...` | A dimension and a measure share the name. Write `@d` or `@m`. | | 422 | `Query::Spec::InvalidSpecError` | `malformed spec: ...` | Unknown keys, a bad predicate name, a bad decorator type or a segment that is not well formed. See [The query spec](#the-query-spec). | | 422 | `ActiveRecord::RecordInvalid` | `Validation failed: ...` | A calculation or filter failed validation: a raw column in a formula, an unresolved `[Name]@m` reference, a blank filter value. | | 422 | `Planner::ResolutionError` | `At least one projection required` | Project at least one field. | | 422 | `Planner::ResolutionError` | `No universe can resolve the query within datasource .` | The fields cannot be joined. See [Query resolution](#query-resolution). | | 422 | `Planner::ResolutionError` | `Could not find path for ` / `Could not find Universe for measure\|dimension: ...` | Same cause, named more precisely. | | 422 | `Planner::ResolutionError` | `cannot parse date ...` | Use `YYYY-MM-DD` or a relative date such as `28d`, `3m`, `1y`. See [Filters](#filters). | | 422 | `Planner::ResolutionError` | `Unsupported filter predicate: ...` | The predicate does not apply to that field's data type. | | 422 | `Planner::ResolutionError` | `Calculation references itself through ` | Break the cycle between calculations. | | 422 | `Planner::ResolutionError` | `Could not resolve segment : ...` | The segment's keys are not reachable from the measures it constrains. See [Segments](#segments). | | 422 | `Planner::ResolutionError` | `
is not a snapshot table` | A snapshot measure on a table without `snapshot:`. | | 422 | `Planner::SecurityPolicyError` | `Security policy context dimension '' is not reachable in the universe for this query. Cannot safely enforce security policy.` | See [Query resolution](#query-resolution). | | 422 | `Planner::SecurityPolicyError` | `Security policy context dimension '' has a composite expression and cannot be used as a security context dimension.` | Point the policy at a plain column dimension. | | 422 | `Unimplemented` | `not implemented: ...` | The combination is not supported yet (a custom predicate, a calculation over a rule-bearing measure). | `zsql explain` is the fastest way to see why: it prints the node tree, which table each node reads and the paths it joined. See [Explain](#explain) and the full list in [Query errors](#errors-and-corrections). ## Query resolution ### `No universe can resolve the query within datasource .` The request asked for a measure grouped or filtered by a dimension its fact cannot reach, and no blend between facts covers it. 0sql refuses rather than guess a join. Check, in order: 1. **Is there a relation?** The dimension's table must be reachable from the measure's fact through `many_to_one` (or `one_to_one`) joins. Add the missing `rel.*.yml`. 2. **Is the change deployed?** `zsql deploy`, then `zsql tables` and `zsql fields ` show what the branch has. 3. **Is it the wrong concept?** If the dimension exists under a different name on the fact's side (`Ship Country` vs `Country`), it is not the same field. Unify the names if they are one concept, or ask for the right one. 4. **Should the measure ignore it?** If the measure is meant to be grouped only by some dimensions, model that with an [exclusion](#exclusions). Remove fields one at a time to find the culprit, or ask `zsql explore --expr ""` which dimensions and measures can still be added. ### Every plan runs in one datasource A query is answered from exactly one datasource; results are never merged across them. If the fields you need live in two, define the missing measure on a table in the other datasource under the same name, and routing picks whichever datasource can serve everything. See [Semantic routing](#semantic-routing). ### `Security policy context dimension '' is not reachable in the universe for this query.` A policy fired (the query projects a tagged or named field) but the node's universe has no path to the policy's `context_dimension`. The query is refused because the filter or mask could not be applied. Either give that table a join path to the context dimension, or pick a context dimension every table with a triggered field can reach. See [Security](#row-level-security). ## Fuzzy field corrections A field reference in a spec or an `expr` is matched by name, uid or synonym, case-insensitively. When nothing matches exactly, the planner scores every field of the wanted kind by trigram overlap with the reference and applies this rule: - A reference shorter than four characters never corrects; it resolves exactly or fails. - If the best candidate scores at least 0.5 and leads the runner-up by at least 0.2, it is used silently and reported in the response's `corrections` array. - Otherwise the request fails with `Semantic::NotFound`, with up to three candidates scoring at least 0.25 appended as ` Did you mean: A, B, C?`. ```json {"sql": "...", "corrections": [{"term": "departmnt", "field_uid": "department", "field_name": "Department", "score": 0.8}]} ``` `corrections` is omitted when empty. `zsql sql` prints each one to stderr as `departmnt → Department (~0.8)`. If a correction surprises you, use the exact name or the uid, and consider adding the misspelling as a `synonym` on the field so it resolves exactly. ## Auth errors | Status | Class | Message | Fix | |---|---|---|---| | 401 | `Unauthorized` | `an API key is required: Authorization: Bearer ` | Send the header. `zsql` reads the key from `ZSQL_API_KEY`, then `.zsql`, then `~/.zsql/config`; run `zsql auth --api-key ...` in the project. | | 401 | `Unauthorized` | `this API key is not valid` | Revoked, rotated or mistyped. Make a new one in the console. | | 403 | `Forbidden` | `query keys are read only; deploy with your personal key` | Use a `zsk_` key for `deploy`, `check`, `test` and `remove`. | | 403 | `Forbidden` | `query key is not granted /` | An owner grants the key that project, with `branch` as a name, `*`, or omitted for the production branch. | | 403 | `Forbidden` | ` has access to ; this needs ` | Ask an owner to raise your level. Read: query. Write: deploy. Owner: access and settings. | | 403 | `Forbidden` | ` has no access to ` | The project is `restricted`; an owner adds you. | | 403 | `Forbidden` | `deploying to the protected production branch '' of needs an owner` | Deploy a feature branch, or have an owner deploy. | See [Accounts, keys and access](#accounts-keys-and-access). ## Numbers look wrong 0sql returns SQL, so a wrong number is a wrong model or a wrong request. `zsql explain` shows which table and joins produced it. ### Numbers are too high (double counting) - **Cardinality is wrong.** A relation declared `many_to_one` that is really one-to-many multiplies the measure. Check the data and fix `cardinality`. See [Cardinality](#cardinality). - **`allow_measure_expansion: true` on a join that is not safe.** Expansion lets a measure be grouped by dimensions on the many side. Remove it unless the join really is safe. - **Many-to-many modelled as a direct join.** Use a junction table with two relations. ### The same measure gives different numbers in two requests When a measure is defined on more than one table, routing chooses per request by the requested dimensions, then by `cost`. If two definitions of `Store Net Paid` disagree, they are not the same concept or one has a bug. Run `zsql explain` on both requests to see which table served each; give a distinct name to anything that means something different. ### A balance or inventory total is wrong across months Summing a daily balance over a month is meaningless. Model it as a [snapshot measure](#snapshot-measures) with `snapshot: ending` (or `beginning`) on a table that declares its `snapshot:` date dimension. ### A cross-domain compound measure repeats a value across rows Expected. When a request groups by a dimension only some components of a compound measure can reach, the others are auto-levelled: aggregated without that dimension and repeated across it. For a rate this is the fixed denominator you want. See [Compound measures](#compound-measures). ## Getting help 1. `zsql check` for the full list of loader errors and warnings. 2. `zsql explain --json --expr "..."` for the plan of a refused or surprising request. 3. Narrow it down: `git diff HEAD~5 -- models/` shows what changed; drop fields until the request plans. 4. Send the `Error in` lines, the explain JSON, the relevant YAML and what you expected. ## Next steps - [Query errors](#errors-and-corrections) - [Deploy](#deploying) and [Tests](#tests) - [Accounts, keys and access](#accounts-keys-and-access) --- ## Canonical YAML Examples These are minimal, valid examples for each 0sql file type, plus one request. ### Table Example (tbl.orders.yml) ```yaml name: Orders physical_name: orders datasource: warehouse cost: 100 fields: - type: dimension name: Order ID data_type: integer expression: primary_key: true sql: order_id - type: dimension name: Customer ID data_type: integer expression: sql: customer_id - type: dimension name: Order Date data_type: date expression: sql: order_date - type: measure name: Order Count data_type: integer expression: sql: count(*) - type: measure name: Total Revenue data_type: decimal expression: sql: sum(order_total) ``` ### Relation Example (rel.sales.yml) ```yaml datasource: warehouse orders_customers: left: Orders right: Customers sql: left.customer_id = right.customer_id cardinality: many_to_one orders_products: left: Order Items right: Products sql: left.product_id = right.product_id cardinality: many_to_one ``` ### Project Example (project.yml) ```yaml # The uid is what the service knows the project as; keep it stable. name: My Analytics Project uid: my-analytics-project production_branch: main ``` ### Datasources Example (datasources.yml) ```yaml # adapter and tier are what the planner uses. Connection keys are carried as # metadata for your own application; 0sql never uses them, and secrets are # stripped before deploy. warehouse: name: Warehouse adapter: postgres tier: hot host: localhost port: 5432 database: analytics schema: public username: analyst ``` ### Request Example Deploy with `zsql deploy`, then POST a spec and a context: ```http POST https://app.0sql.io/projects/my-analytics-project/branches/main/sql Authorization: Bearer zqk_... Content-Type: application/json ``` ```json { "spec": { "projections": [ {"field": "Order Date", "decorators": [{"type": "truncate", "grain": "month"}]}, {"field": "Total Revenue"}, {"field": "Order Count"} ], "filters": [ {"field": "Order Date", "predicate": "greater_than_or_equal_to", "value": "12m"} ], "calculations": [ {"alias": "Revenue per Order", "sql": "[Total Revenue]@m / nullif([Order Count]@m, 0)", "data_type": "decimal"} ] }, "context": {"email": "ana@example.com", "groups": [{"name": "Sales", "tags": ["region:emea"]}]} } ``` The same request as one shorthand line: ```sh zsql sql --expr "month(order date), total revenue, order count, order date >= 12m, Revenue per Order := [Total Revenue]@m / nullif([Order Count]@m, 0)" ``` --- ## Common Mistakes to Avoid ### Wrong: One Name for Two Different Concepts ```yaml # In tbl.calls.yml - type: dimension name: Country # the caller's country # In tbl.shipments.yml - type: dimension name: Country # the ship-to country: WRONG, 0sql merges both into one "Country" ``` **Fix:** Different concepts get different names: "Caller Country" and "Ship Country". The same name on two tables is only correct when it is the same concept (e.g. "Total Revenue" on store and catalog sales), in which case the planner picks the table by dimensions and cost. ### Wrong: Many-to-Many Relationship ```yaml users_roles: cardinality: many_to_many # ERROR: Not supported ``` **Fix:** Create a junction table (user_roles) with two relationships ### Wrong: Measure Without Aggregation ```yaml - type: measure name: Revenue expression: sql: amount # ERROR: No aggregation function ``` **Fix:** Use `sql: sum(amount)` ### Wrong: Dimension With Aggregation ```yaml - type: dimension name: Customer Name expression: sql: max(customer_name) # WRONG: a dimension should not aggregate ``` **Fix:** Use `sql: customer_name` ### Wrong: Missing Required Fields ```yaml name: Orders physical_name: orders # ERROR: Missing datasource, cost, and fields ``` **Fix:** Include all required fields: datasource, name, physical_name, cost, fields ### Wrong: Using Web Links in Field References ```yaml expression: sql: [Total Revenue] - [/semantic-model/fields/cost] # ERROR: URL instead of field name ``` **Fix:** Use field names only: `[Total Revenue] - [Total Cost]` ### Wrong: Spec Without a Context on a Branch With Policies ```json {"spec": {"projections": [{"field": "Customer Email"}, {"field": "Total Revenue"}]}} ``` The branch's `security.yml` has policies, so the service answers `400 ContextRequired: this branch has security policies; a context is required to plan`. **Fix:** Send the caller's security context with every request: `{"spec": {...}, "context": {"email": "ana@example.com", "groups": [{"name": "Sales", "tags": ["region:emea"]}]}}` ### Wrong: Bare Column in a Calculation ```json {"alias": "Margin", "sql": "(order_total - cost) / order_total"} ``` Calculations are written over semantic fields, never warehouse columns. The error is `Validation failed: Sql columns like order_total,cost are not permitted in calculations.` **Fix:** Reference measures and dimensions by name: `"sql": "([Total Revenue]@m - [Total Cost]@m) / nullif([Total Revenue]@m, 0)"`