# CalCOFI integrated database release v2026.09.05

**Release date:** 2026-09-05

## Every dataset has a record: `datasets.json` (the dataset catalog, Phase 0)

The release now writes **`datasets.json`** beside `catalog.json` — one generated record per
`dataset_key` (schema 1.0; `calcofi4db::build_dataset_catalog()`, ≥ 4.1.0) joining what the release
already measured (the `metadata.json` dataset block, `coverage.json` rolled up per dataset —
years, stations, variables, taxa, depth span, the env variables a dataset contributes to another
category — and the content-addressed `catalog.json` objects that belong to it) with the reviewable
registries and with what the live services answer at release time. Each record carries
`distributions[]` (every endpoint: parquet objects with bytes/sha256/since, the CF netCDF, the
ERDDAP ids that exist on erddap.calcofi.io, the ISO 19115 record, the ingest notebook, the
calcofi.org page, the source portal, and the curated mirrors/archives — CoastWatch, EDI, NCEI,
OBIS, the IPT — with `status` and `superseded_by`), `registrations[]` (per portal
`published | planned | n/a`; ERDDAP and OBIS measured, Zenodo from the release DOI), `status`
(stage, priority, issue, blockers, open questions) and `visibility` (`public | internal`). It also
lists `holdings[]` (datasets CalCOFI has but has not ingested, from a sidecar with
`status: planned | external | archived`) and `reference[]` (cruise, ship, grid, spatial, the 19
boundary layers, the GEBCO bathymetry). One `datasets/{dataset_key}.json` per dataset sits beside
it. Nothing on calcofi.io's dataset pages (Phase 1) is authored by hand: they read this file.

- **A release gate**: `check_dataset_catalog()` fails the release on a record without a name, a
  registered category and provider, a description, a bbox or a download; a missing citation is
  exempt only while a provider question covers it (the citation contract's rule); every listed URL
  must answer a one-byte ranged GET (behind `CALCOFI_SKIP_LINK_CHECK`); `datasets.json` joins
  `RELEASE_REQUIRED_OBJECTS`, so `promote_release()` refuses a release without it, and
  `test_release.qmd` checks the file against `datasets.schema.json`, counts it against
  `metadata.json` and re-runs the check before promoting. At v2026.09.04 the finding table is
  `no_citation` × 5 (zoodb, zooscan, farallon, pic-zooplankton, cufes — all exempt, questions open).
- **Three new registries** under `metadata/`: `distribution.csv` (27 curated endpoints —
  the OBIS dataset `0e223f55…` and its IPT resource `calcofi_ichthyo`, eight CoastWatch mirrors of
  the ichthyoplankton, the SIO hydrographic mirrors, EDI/NCEI/DataZoo records, and the seven legacy
  erddap.calcofi.io ids marked `superseded` with their successor), `portal.csv` (16 portals with
  `harvests_from_us` and `observe_method`) and the generated `holdings.csv`;
  `dataset_status.csv` gains `publish_ncei` and `publish_caloos`, and `publish_erddap` now says
  `done` for the 16 datasets erddap.calcofi.io serves.
- **Descriptive metadata leaves the notebooks** (plan § D-9): a dataset's citation, licence, DOI,
  links, contact, keywords, creators and narrative now live in
  `metadata/{provider}/{dataset}/dataset_meta.yml`, the file a provider edits through the
  `metadata` tab of their question Sheet; the notebook keeps the structural keys. `read_calcofi_meta()`
  merges the two, so the release `dataset` table and every consumer see exactly what they saw
  before; a descriptive key left in a notebook now fails the workflows index.
- `coverage.json` `datasets[]` gains `life_stages` per dataset (the dataset's own values).
- Imported the CalOOS working sheet (41 rows) via the new idempotent `scripts/import_caloos_sheet.R`: 24 rows
  matched to already-integrated datasets became `dataset_meta.proposed.yml` proposals (creators, contact,
  keywords, funding, associated parties, QC notes, maintenance) plus 5 new `distribution.csv` rows (3 CalOOS
  module ids, a DataZoo phytoplankton mirror, a NOAA seabird/mammal transect-effort source); 17 unmatched rows
  became new holding sidecars (`metadata/{provider}/{dataset}/dataset_meta.yml`), including the discovery that
  EDI package knb-lter-cce.104 is mislabeled in the sheet (titled "nitrate isotopes", actually POC/PON). Added
  providers `jcvi`, `calpoly`, `stanford`; added category *Genomics & eDNA* and widened *Nutrients & Chemistry*
  / *Phytoplankton*. Filled GCMD Science Keywords (`keywords_gcmd`, 2–5 each, verified against the live GCMD
  KMS export) for all 16 ingested datasets.
- Descriptive dataset metadata split out of the 16 ingest notebooks into per-dataset sidecars
  (`metadata/{provider}/{dataset}/dataset_meta.yml`), editable by providers in a new `metadata` tab of
  their Google Sheet; `scripts/migrate_dataset_meta.R` did the one-off move byte-identically (117 keys,
  comments preserved, the release `dataset` table unchanged before/after), `scripts/sync_dataset_meta_sheets.R`
  does the push/pull (a `holdings` tab in the `calcofi` Sheet is the triage board for the 17 holdings). Both sheet
  scripts now authenticate as the calcofi-admin service account only (`scripts/lib_google_auth.R`), never interactively.

## Contents (generated)

| table | rows | |
|---|---:|---|
| `climatology` | 768,880 | partitioned |
| `cruise` | 842 |  |
| `dataset` | 16 |  |
| `dataset_taxon` | 1,917 |  |
| `grid` | 218 |  |
| `lookup` | 26 |  |
| `measurement_type` | 200 |  |
| `obs` | 26,265,248 | partitioned |
| `obs_attribute` | 458,184 |  |
| `obs_bio` | 1,258,665 |  |
| `obs_env` | 25,006,583 | partitioned |
| `region` | 4 |  |
| `sample` | 1,469,155 |  |
| `sample_measurement` | 589,603 |  |
| `sample_spatial` | 929,664 |  |
| `ship` | 49 |  |
| `spatial` | 13,206 |  |
| `spatial_attribute` | 148,461 |  |
| `taxon` | 2,614 |  |
| `taxon_group` | 441 |  |
| `obs_ctd_full` | 271,394,164 | supplemental |
| `obs_mets_full` | 19,927,416 | supplemental |
| `sample_root` | 421,454 | supplemental |

**23 tables, 348,657,010 rows, ? GB.**

**Datasets (16):** `calcofi_bottle`, `calcofi_ctd-cast`, `calcofi_dic`, `calcofi_mets`, `calcofi_phyllosoma`, `calcofi_phytoplankton`, `cce-lter_euphausiids`, `cce-lter_picoplankton-bacteria`, `cce-lter_zoodb`, `cce-lter_zooscan`, `cdfw_dungeness-crab`, `farallon_bird-mammal`, `sio_mesopelagic-fish`, `sio_pic-zooplankton`, `swfsc_cufes`, `swfsc_ichthyo`

**Software:** calcofi4db 4.1.0, calcofi4r 1.18.0, calcofi4py 0.6.0.

## How to cite

> CalCOFI (2026). CalCOFI Integrated Database, release v2026.09.05 [Data set]. Scripps Institution of Oceanography, NOAA Fisheries, and California Department of Fish and Wildlife. https://calcofi.io/db-schema/?v=v2026.09.05

Cite the source datasets you use alongside the release:

- `calcofi_bottle` — CalCOFI. (2023). CalCOFI Bottle Database 194903-202105. CalCOFI.org. · *license pending*
- `calcofi_ctd-cast` — CalCOFI. (2023). CalCOFI CTD Cast Files. CalCOFI.org. · *license pending*
- `calcofi_dic` — Keeling, C.D.; Lueker, T.J.; Emanuele, G.; Dickson, A.G.; Martz, T.R.; Wolfe, W.H.; Mau, A. (2025). Discrete profile dissolved inorganic carbon, total alkalinity, water temperature and salinity measurements for CalCOFI (NCEI Accession 0301029). NOAA NCEI. https://doi.org/10.25921/3w9f-jd72 · CC-BY-4.0
- `calcofi_mets` — CalCOFI. Underway (METS) TSG/Meteorology Data. CalCOFI.org. · *license pending*
- `calcofi_phyllosoma` — CalCOFI - Scripps Institution of Oceanography and T. Koslow. 2017. Data pertaining to lobster phyllosoma, Panulirus interruptus, collection methods, locations, identification and staging (1951-2008, months of July and August) ver 4. Environmental Data Initiative. https://doi.org/10.6073/pasta/9e38121ebb26f1b59b7b39b2eff844fa · custom (https://portal.edirepository.org/nis/metadataviewer?packageid=knb-lter-cce.188.4)
- `calcofi_phytoplankton` — CalCOFI - Scripps Institution of Oceanography, California Current Ecosystem LTER, and E. Venrick. 2023. Temporal and spatial changes of the abundance and species composition of phytoplankton in the California Current from samples collected aboard CalCOFI cruises from summer 1996 through 2022. ver 4. Environmental Data Initiative. https://doi.org/10.6073/pasta/60edabfbfd85c623fce05822befaa071 · CC0-1.0 (https://creativecommons.org/publicdomain/zero/1.0/)
- `cce-lter_euphausiids` — Ohman, M.D. 2022. California Current Ecosystem Euphausiid data, Brinton and Townsend Euphausiid Database (BTEDB) ver 1. Environmental Data Initiative. https://doi.org/10.6073/pasta/4a92a0044bcd1523a4f994ece874a57d · custom (https://portal.edirepository.org/nis/metadataviewer?packageid=knb-lter-cce.313.1)
- `cce-lter_picoplankton-bacteria` — Landry, M. (2004-2023). Picoplankton and Bacteria Abundance (CalCOFI Cruise). CCE LTER. · *license pending*
- `cce-lter_zoodb` — *citation pending* · custom (https://oceaninformatics.ucsd.edu/zoodb/)
- `cce-lter_zooscan` — *citation pending* · custom (https://oceaninformatics.ucsd.edu/zooscandb/)
- `cdfw_dungeness-crab` — Rogers-Bennett, L.; Jones, E.; Klemmedson, A. (2026). CDFW Dungeness Crab Megalopae from archived CalCOFI plankton samples (1949-2014). California Department of Fish and Wildlife, published through CalCOFI / Scripps Institution of Oceanography.
 · CC-BY-4.0
- `farallon_bird-mammal` — *citation pending* · custom (https://oceanview.pfeg.noaa.gov/CalCOFI/app/resources/docs/Data_Sharing_Agreement_FarallonInstitute.pdf)
- `sio_mesopelagic-fish` — Koslow, J. Anthony (2016). CalCOFI Trawl Data. In California Cooperative Oceanic Fisheries Investigations (CalCOFI): Acoustic and Trawl Data. UC San Diego Library Digital Collections. https://doi.org/10.6075/J0BZ64DH
 · CC-BY-4.0
- `sio_pic-zooplankton` — *citation pending* · *license pending*
- `swfsc_cufes` — *citation pending* · custom (https://coastwatch.pfeg.noaa.gov/erddap/tabledap/erdCalCOFIcufes.das)
- `swfsc_ichthyo` — NOAA Fisheries SWFSC. CalCOFI Ichthyoplankton Database. · *license pending*

## Access

```r
con <- calcofi4r::cc_get_db(version = "v2026.09.05")
```
```python
con = calcofi4py.cc_get_db("v2026.09.05")
```
Parquet: `https://storage.googleapis.com/calcofi-db/ducklake/releases/v2026.09.05/parquet/{table}.parquet`; 
full history: [RELEASES.md](https://storage.googleapis.com/calcofi-db/ducklake/releases/RELEASES.md).
