Skip to the page
Get started

Docs · Concepts

Data sources

One query against one connection — the types available, running one across many machines, and what a run keeps.

A data source is one query against one connection. It belongs to exactly one Check. A Check with one source can only be a flat report; a comparison needs at least two.

#What you can connect to

Sixteen drivers, grouped in the add menu the way you would look for them.

Services

APIREST, with parameters, auth, headers, body and an extract step

Files

FileCSV, JSON, XML, Excel

Databases

Default port
PostgreSQL5432
MySQL3306
MariaDB3306
Microsoft SQL Server1433
Azure SQL Database1433
Oracle1521
SQLite—
CockroachDB26257

Warehouses

Default port
Amazon Redshift5439
Amazon Aurora MySQL3306
Greenplum5432
ClickHouse8123
Vertica5433
Snowflake—

Not relational

Default port
MongoDB27017

Several of these share a driver with another product — MariaDB runs on the MySQL driver, Redshift and Greenplum on the PostgreSQL one — and the menu says so on the entry rather than hiding it.

The catalogue also lists eight products it cannot reach, each with the reason: Apache Cassandra, IBM Db2 LUW, H2, HSQLDB, Apache Derby, Exasol, Sybase ASE and Apache Hive. They are shown rather than omitted so you do not go looking.

#The editor

What you see depends on the driver.

DriverTabs
APIParams · Auth · Headers · Body · Extract · Settings
SQLConnection · Query · Binds · Settings
FileSource

Normalise and Result are added where they apply.

#Running one against many machines

A SQL source can run the same query against more machines. The panel is called Collection, and each entry replaces the server named above it — port, database, credentials and query stay shared. Rows come back carrying a _source column naming the host they came from.

There are three places that list can come from:

  • Own list — kept on the source itself.
  • Library collection — shared. Edit the list once and every data source using it follows.
  • Business Locations — the hosts are read from a property on your places, re-read at every run.

Oracle needs an Easy Connect string for this, not a TNS alias.

Each host reports its own outcome — answered, unreachable or failed, with a row count — and the run log has a Hosts tab showing them. A probe reports partial success honestly: "Connected in 340 ms — 2 of 20 hosts unreachable".

#What happens when a source runs

Four steps, in this order:

  1. Parameter values are substituted, including any recode for this particular source.
  2. The driver runs the query, fanning out across the collection if there is one.
  3. Normalise transforms are applied — after the fan-out, so a twenty-host union is normalised as one result rather than per member.
  4. Row filters remove excluded rows, and the run always says how many it left out.

#What is kept, and what is not

Kept: when it ran, whether it succeeded, how many rows, how long it took, any error — and the shape of what came back, so aggregations know which columns exist.

Not kept: the rows. They live in memory, capped across all sources, and are deliberately not persisted — rows can be large, they go stale, and the source of truth is always the upstream system.

This is why a source has to be run once before the column pickers downstream have anything to offer, and why they empty again after a restart until something runs.

The result pane offers Table, Raw, Schema, Headers, Request, Query and Pipeline, depending on the driver. The Request tab is worth knowing: it shows what Mosaic sent, captured before the call, with credentials redacted — so a connection that failed outright still tells you what it tried.

#The library

A data source can be kept in the library, outside any Check: "Connections and queries kept outside any check. Adding one to a check copies it — what the check holds is then its own."

That copy is the important part. Editing a library source does not change the Checks that took a copy of it. Collections and variables behave the other way round — they stay linked.