Dataset Schema Framework
Use conditional-logic schema generation to keep Dataset markup aligned with real page state and avoid stale or invalid static templates.
When To Use It
Google Search CentralDataset documentationAdd Dataset markup to the canonical landing page for each dataset. Results surface in Google Dataset Search, not as a standard Google Search rich result, so success means the dataset is discoverable and correctly described there.
Key Implementation Documentation Highlights
- Where it shows: datasets appear in Google Dataset Search, not as a rich result in standard Google Search. Mark up the canonical landing page of each dataset.
- What counts as a dataset: a table or CSV file, an organized collection of tables, a proprietary data file, a collection of files that together form a meaningful dataset, a structured object in another format, images capturing data, or files relating to machine learning.
- Required properties: name and description. The description is a summary between 50 and 5000 characters; the name is descriptive (for example "Snow depth in the Northern Hemisphere") and unique for distinct datasets whenever possible.
- Recommended properties: alternateName, creator, citation, funder, hasPart, isPartOf, identifier, isAccessibleForFree, keywords, license, measurementTechnique, sameAs, spatialCoverage, temporalCoverage, variableMeasured, version, url, includedInDataCatalog, distribution.
- creator and funder: a Person (name, givenName, familyName, sameAs) or an Organization (name, url, sameAs). Identify people with an ORCID ID and institutions with a ROR ID in sameAs.
- identifier: a DOI or Compact Identifier, repeated when the dataset has more than one. citation lists academic articles the provider recommends citing in addition to the dataset; don't use it for the dataset's own citation, use identifier for that.
- distribution: DataDownload objects with a required contentUrl (the download link) and a recommended encodingFormat (the file format).
- temporalCoverage: ISO 8601 date or interval, such as
2008,1950-01-01/2013-12-18, or2013-12-19/..for an open-ended interval. - spatialCoverage: a named place as text, a GeoCoordinates point, or a GeoShape box, circle, line or polygon. Points are space separated latitude longitude pairs, latitude first; a box is
lat1 long1 lat2 long2. - license: a URL that unambiguously identifies the specific license version. Collections use hasPart on the parent and isPartOf on the smaller datasets.
- Republished datasets: point sameAs at the canonical page; use isBasedOn when the copy has significant metadata changes or is derived or aggregated from other sources.
- Google expects content parity: every marked-up value must match what users can see on the landing page. Validate with the Rich Results Test and allow several days for Google to crawl and index new pages.
Template Approach
const pageData = {
title: "Charles River Hourly Water Temperature, 2016 to 2025",
summary: "Hourly water temperature readings from 12 monitoring stations along the Charles River, Massachusetts, collected from 2016 to 2025.",
pageUrl: "https://www.example.com/data/charles-river-water-temperature",
licenseUrl: "https://creativecommons.org/licenses/by/4.0/",
publisher: "Example River Institute"
};
const staticTemplate = {
"@context": "https://schema.org",
"@type": "Dataset",
"name": pageData.title,
"description": pageData.summary,
"url": pageData.pageUrl,
"license": pageData.licenseUrl,
"creator": {
"@type": "Organization",
"name": pageData.publisher
}
};
{
"@context": "https://schema.org",
"@type": "Dataset",
"name": "Charles River Hourly Water Temperature, 2016 to 2025",
"description": "Hourly water temperature readings from 12 monitoring stations along the Charles River, Massachusetts, collected from 2016 to 2025.",
"url": "https://www.example.com/data/charles-river-water-temperature",
"license": "https://creativecommons.org/licenses/by/4.0/",
"creator": {
"@type": "Organization",
"name": "Example River Institute"
}
}
Conditional-logic Framework
const ISO_DATE_TIME = /^(\d{4})-(\d{2})-(\d{2})(?:T(\d{2}):(\d{2})(?::(\d{2}))?(Z|[+-]\d{2}:\d{2})?)?$/;
function isAbsoluteUrl(value) {
return typeof value === "string" && /^https?:\/\/[^\s/?#]+\.[^\s/?#]+(?:[/?#]\S*)?$/i.test(value);
}
// Returns the ISO 8601 string unchanged (timezone offset kept), or null if unparseable.
function toIsoDateTime(value) {
if (typeof value !== "string") return null;
const text = value.trim();
const m = text.match(ISO_DATE_TIME);
if (!m) return null;
const [, y, mo, d, h = "00", mi = "00", s = "00"] = m;
const date = new Date(Date.UTC(+y, +mo - 1, +d));
if (date.getUTCMonth() !== +mo - 1 || date.getUTCDate() !== +d) return null;
if (+h > 23 || +mi > 59 || +s > 59) return null;
return text;
}
function toIsoDate(value) {
const iso = toIsoDateTime(value);
return iso ? iso.slice(0, 10) : null;
}
// Drops null, undefined, "", [] and {} recursively.
function compact(value) {
if (Array.isArray(value)) {
const items = value.map(compact).filter((v) => v !== undefined);
return items.length ? items : undefined;
}
if (value && typeof value === "object") {
const out = {};
for (const [key, v] of Object.entries(value)) {
const c = compact(v);
if (c !== undefined) out[key] = c;
}
return Object.keys(out).length ? out : undefined;
}
return value === null || value === undefined || value === "" ? undefined : value;
}
function passesPageGates(source) {
if (source.indexable === false) return false;
if (source.canonicalUrl && source.url && source.canonicalUrl !== source.url) return false;
return source.contentVisible !== false;
}
const ORCID = /^https:\/\/orcid\.org\/\d{4}-\d{4}-\d{4}-\d{3}[\dX]$/;
const ROR = /^https:\/\/ror\.org\/0[a-hj-km-np-tv-z0-9]{6}\d{2}$/;
const DOI = /^10\.\d{4,9}\/\S+$/;
const PARTIAL_DATE = /^(\d{4})(?:-(\d{2})(?:-(\d{2}))?)?$/;
function textLength(value) {
return typeof value === "string" ? value.trim().length : 0;
}
// "2016", "2016-05", "2016-05-01" or a full ISO 8601 date time.
function toTemporalPoint(value) {
const m = typeof value === "string" ? value.match(PARTIAL_DATE) : null;
if (m && m[3]) return toIsoDate(value);
if (m) return !m[2] || (+m[2] >= 1 && +m[2] <= 12) ? value : null;
return value && value.includes("T") ? toIsoDateTime(value) : null;
}
// ISO 8601 date or interval: "2008", "1950-01-01/2013-12-18", or "2013-12-19/.." for an open end.
function toTemporalCoverage(value) {
if (typeof value !== "string") return null;
const [start, end, extra] = value.trim().split("/");
if (extra !== undefined || !toTemporalPoint(start)) return null;
if (end === undefined) return start;
if (end === "..") return `${start}/..`;
if (!toTemporalPoint(end)) return null;
const first = start.length === 4 ? `${start}-01-01` : start.length === 7 ? `${start}-01` : start.slice(0, 10);
const last = end.length === 4 ? `${end}-12-31` : end.length === 7 ? `${end}-31` : end.slice(0, 10);
return first <= last ? `${start}/${end}` : null; // the interval must not run backwards
}
function isLatLong(lat, long) {
return typeof lat === "number" && typeof long === "number" && Math.abs(lat) <= 90 && Math.abs(long) <= 180;
}
// A named place, a GeoCoordinates point, or a GeoShape box ("lat1 long1 lat2 long2", latitude first).
function buildSpatialCoverage(place) {
if (typeof place === "string") return place;
if (!place) return null;
if (place.point) {
const { lat, long } = place.point;
if (!isLatLong(lat, long)) return null;
return { "@type": "Place", name: place.name, geo: { "@type": "GeoCoordinates", latitude: lat, longitude: long } };
}
if (Array.isArray(place.box) && place.box.length === 4) {
const [lat1, long1, lat2, long2] = place.box;
if (!isLatLong(lat1, long1) || !isLatLong(lat2, long2)) return null;
return { "@type": "Place", name: place.name, geo: { "@type": "GeoShape", box: place.box.join(" ") } };
}
return place.name || null;
}
// Person: name, givenName, familyName, url, ORCID sameAs. Organization: name, url, ROR sameAs.
function buildAgent(agent) {
if (!agent || !agent.name) return null;
if (agent.type === "person") {
return {
"@type": "Person",
name: agent.name,
givenName: agent.givenName,
familyName: agent.familyName,
url: isAbsoluteUrl(agent.url) ? agent.url : null,
sameAs: ORCID.test(agent.orcid || "") ? agent.orcid : null
};
}
return {
"@type": "Organization",
name: agent.name,
url: isAbsoluteUrl(agent.url) ? agent.url : null,
sameAs: ROR.test(agent.ror || "") ? agent.ror : null
};
}
// Bare DOIs become resolvable https://doi.org/ URLs; other identifiers stay as given.
function toIdentifier(value) {
if (typeof value !== "string" || !value.trim()) return null;
const text = value.trim();
return DOI.test(text) ? `https://doi.org/${text}` : text;
}
function buildVariable(variable) {
if (typeof variable === "string") return variable;
if (!variable || !variable.name) return null;
return { "@type": "PropertyValue", name: variable.name, unitText: variable.unit, description: variable.description };
}
function buildDistribution(file) {
// contentUrl is required on every DataDownload.
if (!file || !isAbsoluteUrl(file.contentUrl)) return null;
return { "@type": "DataDownload", contentUrl: file.contentUrl, encodingFormat: file.encodingFormat };
}
// A related dataset as a URL, or as a Dataset with its own name and valid description.
function buildRelatedDataset(part) {
if (isAbsoluteUrl(part)) return part;
const length = textLength(part && part.description);
if (!part || !part.name || length < 50 || length > 5000) return null;
return { "@type": "Dataset", name: part.name, description: part.description.trim(), url: isAbsoluteUrl(part.url) ? part.url : null };
}
function buildDatasetSchema(source) {
// (a) Page gates: the indexable, self-canonical landing page for this dataset, with the metadata visible.
if (!passesPageGates(source) || !isAbsoluteUrl(source.url)) return null;
// (b) Required: name, and a description between 50 and 5000 characters.
const length = textLength(source.description);
if (!source.name || length < 50 || length > 5000) return null;
const identifier = (source.identifiers || []).map(toIdentifier).filter(Boolean);
const isOwnCitation = (c) => identifier.includes(toIdentifier(c));
// (c) Recommended properties, only when the source has them; (d) value rules inline.
const schema = {
"@context": "https://schema.org",
"@type": "Dataset",
"@id": `${source.url}#dataset`,
name: source.name,
description: source.description.trim(),
alternateName: source.alternateNames,
url: source.url,
sameAs: (source.sameAs || []).filter(isAbsoluteUrl), // the canonical page when republished
identifier, // repeat for multiple DOIs or Compact Identifiers
version: source.version,
keywords: source.keywords,
license: isAbsoluteUrl(source.license) ? source.license : null, // a URL for one specific license version
isAccessibleForFree: typeof source.free === "boolean" ? source.free : null,
creator: (source.creators || []).map(buildAgent),
funder: (source.funders || []).map(buildAgent),
citation: (source.citations || []).filter((c) => !isOwnCitation(c)), // related articles, not the dataset itself
temporalCoverage: toTemporalCoverage(source.temporalCoverage),
spatialCoverage: buildSpatialCoverage(source.spatialCoverage),
variableMeasured: (source.variables || []).map(buildVariable),
measurementTechnique: source.measurementTechniques,
isPartOf: (source.partOf || []).map(buildRelatedDataset),
hasPart: (source.parts || []).map(buildRelatedDataset),
includedInDataCatalog: source.catalog && source.catalog.name
? { "@type": "DataCatalog", name: source.catalog.name, url: isAbsoluteUrl(source.catalog.url) ? source.catalog.url : null }
: null,
distribution: (source.files || []).map(buildDistribution)
};
// (e) Strip empty values.
return compact(schema);
}
// Usage:
// const source = {
// "url": "https://www.example.com/data/charles-river-water-temperature",
// "canonicalUrl": "https://www.example.com/data/charles-river-water-temperature",
// "indexable": true,
// "contentVisible": true,
// "name": "Charles River Hourly Water Temperature, 2016 to 2025",
// "description": "Hourly water temperature readings from 12 monitoring stations along the Charles River, Massachusetts, collected from January 2016 to December 2025 and quality controlled in 2026.",
// "alternateNames": [
// "CRWT 2016-2025"
// ],
// "identifiers": [
// "10.5555/example.crwt.2026"
// ],
// "version": "2026.1",
// "keywords": [
// "water temperature",
// "rivers",
// "hydrology",
// "Massachusetts"
// ],
// "license": "https://creativecommons.org/licenses/by/4.0/",
// "free": true,
// "creators": [
// {
// "type": "person",
// "name": "Josiah Carberry",
// "givenName": "Josiah",
// "familyName": "Carberry",
// "orcid": "https://orcid.org/0000-0002-1825-0097",
// "url": "https://www.example.com/people/josiah-carberry"
// }
// ],
// "funders": [
// {
// "type": "organization",
// "name": "Example Watershed Foundation",
// "url": "https://foundation.example.com/",
// "ror": "https://ror.org/0abcdef12"
// }
// ],
// "citations": [
// "https://doi.org/10.5555/example.article.2025"
// ],
// "temporalCoverage": "2016-01-01/2025-12-31",
// "spatialCoverage": {
// "name": "Charles River watershed, Massachusetts",
// "box": [
// 42.1,
// -71.6,
// 42.4,
// -71.0
// ]
// },
// "variables": [
// {
// "name": "Water temperature",
// "unit": "degrees Celsius"
// }
// ],
// "measurementTechniques": [
// "Submerged thermistor loggers, 60 minute sampling interval"
// ],
// "partOf": [
// "https://www.example.com/data/new-england-river-observations"
// ],
// "parts": [
// "https://www.example.com/data/charles-river-water-temperature/2025"
// ],
// "catalog": {
// "name": "Example River Institute Data Catalog",
// "url": "https://www.example.com/data/"
// },
// "files": [
// {
// "contentUrl": "https://www.example.com/data/files/crwt-2016-2025.csv",
// "encodingFormat": "text/csv"
// },
// {
// "contentUrl": "https://www.example.com/data/files/crwt-2016-2025.parquet",
// "encodingFormat": "application/vnd.apache.parquet"
// }
// ],
// "sameAs": [
// "https://catalog.example.org/records/crwt-2016-2025"
// ]
// };
// const schema = buildDatasetSchema(source); // null when a gate or required property fails
import re
from datetime import date, datetime, timezone
ISO_DATE_TIME = re.compile(r"^(\d{4})-(\d{2})-(\d{2})(?:T(\d{2}):(\d{2})(?::(\d{2}))?(Z|[+-]\d{2}:\d{2})?)?$")
ABSOLUTE_URL = re.compile(r"^https?://[^\s/?#]+\.[^\s/?#]+(?:[/?#]\S*)?$", re.I)
def is_absolute_url(value):
return isinstance(value, str) and bool(ABSOLUTE_URL.match(value))
# Returns the ISO 8601 string unchanged (timezone offset kept), or None if unparseable.
def to_iso_date_time(value):
if not isinstance(value, str):
return None
text = value.strip()
m = ISO_DATE_TIME.match(text)
if not m:
return None
y, mo, d = int(m[1]), int(m[2]), int(m[3])
try:
date(y, mo, d)
except ValueError:
return None
if int(m[4] or 0) > 23 or int(m[5] or 0) > 59 or int(m[6] or 0) > 59:
return None
return text
def to_iso_date(value):
iso = to_iso_date_time(value)
return iso[:10] if iso else None
# Drops None, "", [] and {} recursively.
def compact(value):
if isinstance(value, list):
items = [c for c in (compact(v) for v in value) if c is not None]
return items or None
if isinstance(value, dict):
out = {k: c for k, c in ((k, compact(v)) for k, v in value.items()) if c is not None}
return out or None
return None if value is None or value == "" else value
def passes_page_gates(source):
if source.get("indexable") is False:
return False
if source.get("canonicalUrl") and source.get("url") and source["canonicalUrl"] != source["url"]:
return False
return source.get("contentVisible") is not False
ORCID = re.compile(r"^https://orcid\.org/\d{4}-\d{4}-\d{4}-\d{3}[\dX]$")
ROR = re.compile(r"^https://ror\.org/0[a-hj-km-np-tv-z0-9]{6}\d{2}$")
DOI = re.compile(r"^10\.\d{4,9}/\S+$")
PARTIAL_DATE = re.compile(r"^(\d{4})(?:-(\d{2})(?:-(\d{2}))?)?$")
def is_number(value):
return isinstance(value, (int, float)) and not isinstance(value, bool)
# Formats a number the way JavaScript prints it (42.0 becomes "42").
def js_number(value):
return str(int(value)) if float(value).is_integer() else repr(value)
def text_length(value):
return len(value.strip()) if isinstance(value, str) else 0
# "2016", "2016-05", "2016-05-01" or a full ISO 8601 date time.
def to_temporal_point(value):
m = PARTIAL_DATE.match(value) if isinstance(value, str) else None
if m and m[3]:
return to_iso_date(value)
if m:
return value if not m[2] or 1 <= int(m[2]) <= 12 else None
return to_iso_date_time(value) if isinstance(value, str) and "T" in value else None
# ISO 8601 date or interval: "2008", "1950-01-01/2013-12-18", or "2013-12-19/.." for an open end.
def to_temporal_coverage(value):
if not isinstance(value, str):
return None
parts = value.strip().split("/")
start = parts[0]
if len(parts) > 2 or not to_temporal_point(start):
return None
if len(parts) == 1:
return start
end = parts[1]
if end == "..":
return f"{start}/.."
if not to_temporal_point(end):
return None
first = f"{start}-01-01" if len(start) == 4 else f"{start}-01" if len(start) == 7 else start[:10]
last = f"{end}-12-31" if len(end) == 4 else f"{end}-31" if len(end) == 7 else end[:10]
return f"{start}/{end}" if first <= last else None # the interval must not run backwards
def is_lat_long(lat, long):
return is_number(lat) and is_number(long) and abs(lat) <= 90 and abs(long) <= 180
# A named place, a GeoCoordinates point, or a GeoShape box ("lat1 long1 lat2 long2", latitude first).
def build_spatial_coverage(place):
if isinstance(place, str):
return place
if not place:
return None
if place.get("point"):
lat, long = place["point"].get("lat"), place["point"].get("long")
if not is_lat_long(lat, long):
return None
return {"@type": "Place", "name": place.get("name"), "geo": {"@type": "GeoCoordinates", "latitude": lat, "longitude": long}}
box = place.get("box")
if isinstance(box, list) and len(box) == 4:
if not is_lat_long(box[0], box[1]) or not is_lat_long(box[2], box[3]):
return None
return {"@type": "Place", "name": place.get("name"), "geo": {"@type": "GeoShape", "box": " ".join(js_number(n) for n in box)}}
return place.get("name")
# Person: name, givenName, familyName, url, ORCID sameAs. Organization: name, url, ROR sameAs.
def build_agent(agent):
if not agent or not agent.get("name"):
return None
url = agent.get("url") if is_absolute_url(agent.get("url")) else None
if agent.get("type") == "person":
return {
"@type": "Person",
"name": agent["name"],
"givenName": agent.get("givenName"),
"familyName": agent.get("familyName"),
"url": url,
"sameAs": agent.get("orcid") if ORCID.match(agent.get("orcid") or "") else None,
}
return {
"@type": "Organization",
"name": agent["name"],
"url": url,
"sameAs": agent.get("ror") if ROR.match(agent.get("ror") or "") else None,
}
# Bare DOIs become resolvable https://doi.org/ URLs; other identifiers stay as given.
def to_identifier(value):
if not isinstance(value, str) or not value.strip():
return None
text = value.strip()
return f"https://doi.org/{text}" if DOI.match(text) else text
def build_variable(variable):
if isinstance(variable, str):
return variable
if not variable or not variable.get("name"):
return None
return {"@type": "PropertyValue", "name": variable["name"], "unitText": variable.get("unit"), "description": variable.get("description")}
def build_distribution(file):
# contentUrl is required on every DataDownload.
if not file or not is_absolute_url(file.get("contentUrl")):
return None
return {"@type": "DataDownload", "contentUrl": file["contentUrl"], "encodingFormat": file.get("encodingFormat")}
# A related dataset as a URL, or as a Dataset with its own name and valid description.
def build_related_dataset(part):
if is_absolute_url(part):
return part
if not isinstance(part, dict):
return None
length = text_length(part.get("description"))
if not part.get("name") or length < 50 or length > 5000:
return None
url = part.get("url") if is_absolute_url(part.get("url")) else None
return {"@type": "Dataset", "name": part["name"], "description": part["description"].strip(), "url": url}
def build_dataset_schema(source):
# (a) Page gates: the indexable, self-canonical landing page for this dataset, with the metadata visible.
if not passes_page_gates(source) or not is_absolute_url(source.get("url")):
return None
# (b) Required: name, and a description between 50 and 5000 characters.
length = text_length(source.get("description"))
if not source.get("name") or length < 50 or length > 5000:
return None
identifier = [i for i in (to_identifier(v) for v in source.get("identifiers") or []) if i]
catalog = source.get("catalog") or {}
# (c) Recommended properties, only when the source has them; (d) value rules inline.
schema = {
"@context": "https://schema.org",
"@type": "Dataset",
"@id": f"{source['url']}#dataset",
"name": source["name"],
"description": source["description"].strip(),
"alternateName": source.get("alternateNames"),
"url": source["url"],
"sameAs": [u for u in source.get("sameAs") or [] if is_absolute_url(u)], # the canonical page when republished
"identifier": identifier, # repeat for multiple DOIs or Compact Identifiers
"version": source.get("version"),
"keywords": source.get("keywords"),
"license": source.get("license") if is_absolute_url(source.get("license")) else None, # a URL for one specific license version
"isAccessibleForFree": source.get("free") if isinstance(source.get("free"), bool) else None,
"creator": [build_agent(a) for a in source.get("creators") or []],
"funder": [build_agent(a) for a in source.get("funders") or []],
"citation": [c for c in source.get("citations") or [] if to_identifier(c) not in identifier], # related articles, not the dataset itself
"temporalCoverage": to_temporal_coverage(source.get("temporalCoverage")),
"spatialCoverage": build_spatial_coverage(source.get("spatialCoverage")),
"variableMeasured": [build_variable(v) for v in source.get("variables") or []],
"measurementTechnique": source.get("measurementTechniques"),
"isPartOf": [build_related_dataset(p) for p in source.get("partOf") or []],
"hasPart": [build_related_dataset(p) for p in source.get("parts") or []],
"includedInDataCatalog": {
"@type": "DataCatalog",
"name": catalog["name"],
"url": catalog.get("url") if is_absolute_url(catalog.get("url")) else None,
} if catalog.get("name") else None,
"distribution": [build_distribution(f) for f in source.get("files") or []],
}
# (e) Strip empty values.
return compact(schema)
{
"@context": "https://schema.org",
"@type": "Dataset",
"@id": "https://www.example.com/data/charles-river-water-temperature#dataset",
"name": "Charles River Hourly Water Temperature, 2016 to 2025",
"description": "Hourly water temperature readings from 12 monitoring stations along the Charles River, Massachusetts, collected from January 2016 to December 2025 and quality controlled in 2026.",
"alternateName": [
"CRWT 2016-2025"
],
"url": "https://www.example.com/data/charles-river-water-temperature",
"sameAs": [
"https://catalog.example.org/records/crwt-2016-2025"
],
"identifier": [
"https://doi.org/10.5555/example.crwt.2026"
],
"version": "2026.1",
"keywords": [
"water temperature",
"rivers",
"hydrology",
"Massachusetts"
],
"license": "https://creativecommons.org/licenses/by/4.0/",
"isAccessibleForFree": true,
"creator": [
{
"@type": "Person",
"name": "Josiah Carberry",
"givenName": "Josiah",
"familyName": "Carberry",
"url": "https://www.example.com/people/josiah-carberry",
"sameAs": "https://orcid.org/0000-0002-1825-0097"
}
],
"funder": [
{
"@type": "Organization",
"name": "Example Watershed Foundation",
"url": "https://foundation.example.com/",
"sameAs": "https://ror.org/0abcdef12"
}
],
"citation": [
"https://doi.org/10.5555/example.article.2025"
],
"temporalCoverage": "2016-01-01/2025-12-31",
"spatialCoverage": {
"@type": "Place",
"name": "Charles River watershed, Massachusetts",
"geo": {
"@type": "GeoShape",
"box": "42.1 -71.6 42.4 -71"
}
},
"variableMeasured": [
{
"@type": "PropertyValue",
"name": "Water temperature",
"unitText": "degrees Celsius"
}
],
"measurementTechnique": [
"Submerged thermistor loggers, 60 minute sampling interval"
],
"isPartOf": [
"https://www.example.com/data/new-england-river-observations"
],
"hasPart": [
"https://www.example.com/data/charles-river-water-temperature/2025"
],
"includedInDataCatalog": {
"@type": "DataCatalog",
"name": "Example River Institute Data Catalog",
"url": "https://www.example.com/data/"
},
"distribution": [
{
"@type": "DataDownload",
"contentUrl": "https://www.example.com/data/files/crwt-2016-2025.csv",
"encodingFormat": "text/csv"
},
{
"@type": "DataDownload",
"contentUrl": "https://www.example.com/data/files/crwt-2016-2025.parquet",
"encodingFormat": "application/vnd.apache.parquet"
}
]
}
Why Conditional Logic Is Better Than Static Templates
- Returns nothing when the description falls outside 50 to 5000 characters, instead of shipping a record Google rejects.
- Normalizes bare DOIs into resolvable identifier URLs and drops citations that point back at the dataset itself.
- Checks temporalCoverage intervals and GeoShape coordinate order and ranges before they reach the page.
- Drops DataDownload entries without a contentUrl and emits optional fields only when the catalog record has them, so one function scales across every landing page.
Property Reference
What Google's Dataset documentation requires and recommends, property by property. These are the same rules the SchemaCDN validator checks. Each name links to its schema.org definition.
| Property | Expects | Notes | |
|---|---|---|---|
| description | Required | Text | A short summary describing the dataset, between 50 and 5000 characters long. |
| name | Required | Text | A descriptive name of the dataset. Use unique names for distinct datasets. |
| alternateName | Recommended | Text, can repeat | Alternative names that have been used to refer to this dataset, such as aliases or abbreviations. |
| creator | Recommended | Person or Organization, can repeat | The creator or author of this dataset (Person or Organization), ideally identified with ORCID or ROR IDs in sameAs. |
| creator.name | Recommended | Text | |
| creator.sameAs | Recommended | URL, can repeat | An ORCID ID (person) or ROR ID (organization). |
| creator.url | Recommended | URL | For organizations. |
| creator.givenName | Recommended | Text | For persons. |
| creator.familyName | Recommended | Text | For persons. |
| citation | Recommended | Text or CreativeWork, can repeat | Academic articles the data provider recommends be cited in addition to the dataset itself. |
| funder | Recommended | Person or Organization, can repeat | A person or organization that provides financial support for this dataset. |
| hasPart | Recommended | URL or Dataset, can repeat | Use when this dataset is a collection of smaller datasets. |
| isPartOf | Recommended | URL or Dataset, can repeat | Use when this dataset is part of a larger dataset collection. |
| identifier | Recommended | URL, Text or PropertyValue, can repeat | An identifier such as a DOI or Compact Identifier. Repeat for multiple identifiers. |
| isAccessibleForFree | Recommended | Boolean | Whether the dataset is accessible without payment. |
| keywords | Recommended | Text, can repeat | Summarizing the dataset. |
| license | Recommended | URL or CreativeWork | A URL that unambiguously identifies the specific license version under which the dataset is distributed. |
| measurementTechnique | Recommended | Text or URL, can repeat | The technique, technology, or methodology used in the dataset. |
| sameAs | Recommended | URL, can repeat | The URL of a reference web page that unambiguously indicates the dataset's identity. |
| spatialCoverage | Recommended | Text or Place, can repeat | A named place, a GeoCoordinates point, or a GeoShape describing the dataset's spatial aspect. |
| temporalCoverage | Recommended | Text | The time interval the data covers, in ISO 8601 format (for example 1950-01-01/2013-12-18 or 2013-12-19/..). |
| variableMeasured | Recommended | Text or PropertyValue, can repeat | The variable that this dataset measures, for example temperature or pressure. |
| version | Recommended | Text or Number | The version number for the dataset. |
| url | Recommended | URL | Location of a page describing the dataset. |
| includedInDataCatalog | Recommended | DataCatalog | The catalog to which the dataset belongs. |
| distribution | Recommended | DataDownload, can repeat | DataDownload objects describing where to download the dataset and in which format. |
| distribution.contentUrl | Required | URL | The link for the download. |
| distribution.encodingFormat | Recommended | Text or URL | The file format of the distribution. |