Discovering ocean data with AQUAVIEW¶

PAM–Glider Rodeo Hackweek · hands-on tutorial (Python)¶

AQUAVIEW is one common catalog over 600,000+ ocean datasets from around a hundred sources — satellite, ship and float profiles, gliders, buoys, models, biodiversity, vessel traffic, and more. Instead of tracking down data across many separate servers, clouds and portals, you discover and reach it all in one place — by website, by code, or by asking an AI.

This notebook is in two parts:

  1. Meet AQUAVIEW — what it is, and the three ways to reach it.
  2. A worked example — find a real Hawaiʻi glider, pull its track, and check the satellite against it.

By the end you will have reproduced a genuine validation result: satellite SST vs glider SST along a two-week track, with bias, RMSE and correlation.

Everything runs live against the real catalog — no sign-in needed for discovery. There is an identical R version for R users.

What you need: pystac-client, pandas, xarray, netCDF4, matplotlib, requests. On the Glider Rodeo Hub these are already installed.


Part A · Meet AQUAVIEW¶

Meet the tool¶

This is aquaview.org — the front door. One platform, one catalog, and an Explore map, an API, and AI search over it. Today we drive it from code.

aquaview.org homepage — 'One data plane from ingest to access', with the Explore map of global data density

Under the hood it's one catalog underneath, three ways to reach it, ~100 sources feeding in — today we learn the Code door, but the same catalog is behind all three:

AQUAVIEWone common catalog · 600,000+ datasets · ~100 sourcesExplorewebsite · point & clickCodePython · R (STAC API)AIMCP · plain Englishthe same catalog underneathNOAAIOOSArgoOBISmodelsAISbuoyssatellite

Five words we'll use¶

STAC (the standard AQUAVIEW speaks) nests like this — you only need these five:

AQUAVIEW
   └─ Source / Collection      e.g. "PacIOOS"          a data provider or product family
        └─ Dataset / Item      e.g. "ww3_hawaii"       one dataset record, with location + time
             └─ Assets         nc · csv · png · …       the actual downloadable files

A search returns items; each item lists its assets. That's the whole model.

A1 · Connect¶

AQUAVIEW speaks the STAC standard, so the standard client just works — one line, no sign-in for discovery.

One host, defined once. Everything below uses STAC. If AQUAVIEW moves the endpoint, this is the only line you change.

In [1]:
from pystac_client import Client

STAC = "https://service.aquaview.org/stac"
catalog = Client.open(STAC)

print("Connected to AQUAVIEW ✓")
print(catalog.get_self_href())
Connected to AQUAVIEW ✓
https://service.aquaview.org/stac

A2 · One catalog over everything¶

AQUAVIEW federates the whole ocean-data landscape into one searchable place. Here's the size, and a taste of the range.

In [2]:
# One request for the full source list, kept for the rest of the notebook.
collections = list(catalog.get_collections())
total = catalog.search(bbox=[-180, -90, 180, 90]).matched()

print(f"{total:,} datasets   ·   {len(collections)} collections   ·   one catalog, one API\n")

kinds = {
    "Ship & float profiles":        "WOD · Argo (GADR, RG_ARGO) · EN4 · CCHDO",
    "Satellite (SST, colour, SAR)": "CoastWatch · GOES-R · PolarWatch · Cerulean oil-slicks",
    "Gliders & autonomous":         "IOOS Glider DAC · Spray · Voice of the Ocean",
    "Buoys & sensors":              "NDBC · IOOS Sensors · CDIP waves",
    "Ocean & weather models":       "HYCOM · RTOFS · GFS · Copernicus GLORYS",
    "Biodiversity & human activity":"OBIS species · MarineCadastre AIS vessel traffic",
}
for k, v in kinds.items():
    print(f"  {k:<30} {v}")
609,756 datasets   ·   100 collections   ·   one catalog, one API

  Ship & float profiles          WOD · Argo (GADR, RG_ARGO) · EN4 · CCHDO
  Satellite (SST, colour, SAR)   CoastWatch · GOES-R · PolarWatch · Cerulean oil-slicks
  Gliders & autonomous           IOOS Glider DAC · Spray · Voice of the Ocean
  Buoys & sensors                NDBC · IOOS Sensors · CDIP waves
  Ocean & weather models         HYCOM · RTOFS · GFS · Copernicus GLORYS
  Biodiversity & human activity  OBIS species · MarineCadastre AIS vessel traffic

A3 · Find a source without memorising IDs¶

You don't have to know collection IDs in advance. Each collection carries a title, description and keywords; we already have all of them from the single request above, so we can filter locally — no extra calls, and you can see exactly what is being matched.

In [3]:
def haystack(c):
    d = c.to_dict()
    return " ".join([c.id, d.get("title") or "", d.get("description") or "",
                     " ".join(d.get("keywords") or [])]).lower()

HAY = {c.id: haystack(c) for c in collections}

def find_sources(term, n=6):
    """Which sources describe themselves using this word?"""
    return [cid for cid, h in HAY.items() if term.lower() in h][:n]

for term in ["glider", "sea surface temperature", "vessel", "argo"]:
    print(f"  {term:24} → {find_sources(term)}")
  glider                   → ['CALOOS', 'CCHDO', 'CENCOOS', 'IOOS', 'MARACOOS', 'NANOOS']
  sea surface temperature  → ['CDIP', 'COASTWATCH_WC', 'GOES_R', 'NEFSC', 'NOAA_GDP', 'POLARWATCH']
  vessel                   → ['CERULEAN', 'CERULEAN_ARCHIVE', 'CERULEAN_VESSEL_30D', 'MARINECADASTRE_AIS', 'OET_NA176']
  argo                     → ['CCHDO', 'GADR', 'ISAS', 'RG_ARGO']

Read this result carefully — it is the most useful thing in Part A.

This searches how each source describes itself. That is a good first cut, but a source's blurb does not list every dataset inside it. find_sources("wave") misses PacIOOS, even though PacIOOS serves the Hawaiʻi WaveWatch III model we use later — PacIOOS simply doesn't say "wave" in its own description.

So the rule is:

You want… Use
a provider you half-remember find_sources(), above
datasets in your study area a bounding-box search — Part B, and it is the reliable one
to browse visually Explore

When the text search comes up short, search your box instead. That is what B4 does to find the glider.

A4 · Three ways in¶

The same catalog, whichever way you prefer to work — and a result is the same STAC item whichever door you use, so the data looks identical downstream:

Way in For This tutorial
Explore map at aquaview.org/explore click-and-draw discovery, no code shown in B8
STAC API (pystac-client / rstac) reproducible, scriptable the rest of this notebook
AI / MCP chat ask in plain English shown in B8

Part B · A worked example¶

Is the satellite right, along a Hawaiʻi glider track?¶

Your gliders record passive acoustics and CTD off Hawaiʻi. Before you can interpret that, you need the ocean around the track — and you need to know whether the satellite products you plan to lean on actually agree with what the glider measured.

We'll do it in six moves: search a box → narrow → count → find the glider → load its track → match the satellite to it.

B1 · Search one place, one window¶

The core move: a bounding box around the fleet and a date range. One call reaches across every source — no per-server hunting.

Give the datetime as a full RFC 3339 timestamp (2025-01-01T00:00:00Z/..). pystac-client tidies a bare 2025-01-01/.. for you, but the raw API rejects it — so if you ever drop down to requests, use the long form.

In [4]:
HAWAII = [-161, 18, -154, 23]            # west, south, east, north
WHEN   = "2025-01-01T00:00:00Z/.."       # 2025 to now

search = catalog.search(bbox=HAWAII, datetime=WHEN)
print(f"{search.matched():,} datasets match in the Hawaiʻi box (2025–present)")
4,230 datasets match in the Hawaiʻi box (2025–present)

B2 · From a broad catch to what you actually want¶

A wide search returns everything that intersects your box and time — that's discovery working, not failing.

Two things to notice in the table below.

  1. The first page is not a ranking. Results come back in the catalog's own order, not by relevance, so the top rows tell you nothing about the mix. (Cerulean SAR oil-slick detections lead here; they are about a sixth of the box, not the bulk of it.)
  2. Many dates are a range, not a moment. A continuously-updated model or a glider deployment has start_datetime/end_datetime and a null datetime. Reading only datetime shows a blank — so we fall back.
In [5]:
import pandas as pd

def when(p):
    """STAC items are either a moment or a range — show whichever this one has."""
    if p.get("datetime"):
        return p["datetime"][:10]
    s, e = (p.get("start_datetime") or "")[:10], (p.get("end_datetime") or "")[:10]
    return f"{s} → {e}" if s or e else ""

def short(x, n=44):
    """Trim a long title at a word boundary rather than mid-word."""
    x = x or ""
    return x if len(x) <= n else x[:n].rsplit(" ", 1)[0] + "…"

def to_table(item_search, n=6):
    rows = []
    for item in item_search.items():
        p = item.properties
        rows.append({"source": item.collection_id, "id": short(item.id, 26),
                     "when": when(p), "title": short(p.get("title"), 44)})
        if len(rows) >= n:
            break
    return pd.DataFrame(rows)

to_table(search)               # the raw, unfiltered catch
Out[5]:
source id when title
0 CERULEAN 5216945 2026-09-13 Oil slick 5216945
1 CERULEAN 5216944 2026-09-13 Oil slick 5216944
2 CERULEAN 5216943 2026-09-13 Oil slick 5216943
3 CERULEAN 5214248 2026-09-12 Oil slick 5214248
4 CERULEAN 5214247 2026-09-12 Oil slick 5214247
5 CERULEAN 5214246 2026-09-12 Oil slick 5214246

So you narrow — the real discovery skill. Pass collections= to keep only the sources relevant to a glider mission. Same call, one extra argument.

Note the SST source is COASTWATCH, not COASTWATCH_WC. COASTWATCH_WC is the West Coast node — over Hawaiʻi it carries HF-radar currents and ocean colour. The satellite SST products live in plain COASTWATCH, and that is the one we validate against in B6.

In [6]:
relevant = ["COASTWATCH", "PacIOOS", "OBIS", "MARINECADASTRE_AIS", "IOOS"]
focused  = catalog.search(bbox=HAWAII, collections=relevant, datetime=WHEN)
print(f"{focused.matched():,} datasets across the sources we care about\n")

rows = []
for c in relevant:
    item = next(catalog.search(bbox=HAWAII, collections=[c], datetime=WHEN).items(), None)
    if item:
        p = item.properties
        rows.append({"source": c, "id": short(item.id, 26), "when": when(p),
                     "title": short(p.get("title"), 44)})
pd.DataFrame(rows)             # identical columns, whatever the source
227 datasets across the sources we care about

Out[6]:
source id when title
0 COASTWATCH noaacweddymunsterdaily 2020-05-05 → 2026-06-08 MUltiparameter NRT System for Tracking…
1 PacIOOS ww3_hawaii_lon180 2011-06-21 → 2026-09-19 WaveWatch III (WW3) Hawaii Regional Wave…
2 OBIS 345c1b88-2bab-4403-ad2c-17… 2026-06-29 DFO Maritimes Region Cetacean Sightings
3 MARINECADASTRE_AIS mc_dataset_67919 2024-07-01 → 2026-03-01 Vessel Routing Measures
4 IOOS sg626-20250729T1452 2025-07-29 → 2025-10-29 sg626-20250729T1452

B3 · What's here, by source?¶

Before downloading anything, ask how much of each kind exists in your box — one cheap count per source, no data moved. .matched() returns the total without fetching a single file.

In [7]:
import matplotlib.pyplot as plt

of_interest = {
    "COASTWATCH":         "Satellite SST",
    "COASTWATCH_WC":      "Ocean colour / HF radar",
    "PacIOOS":            "Hawaiʻi waves & models",
    "OBIS":               "Species-occurrence context",
    "MARINECADASTRE_AIS": "Vessel-traffic context",
    "IOOS":               "Glider Data Assembly Center",
}
counts = {c: catalog.search(bbox=HAWAII, collections=[c], datetime=WHEN).matched()
          for c in of_interest}

s  = pd.Series(counts).rename(index=of_interest).sort_values()
ax = s.plot.barh(color="#0b7d84", figsize=(8, 3.2))
for i, v in enumerate(s):
    ax.text(v, i, f" {v:,}", va="center", fontsize=9, color="#334")
ax.set_title("Datasets in the Hawaiʻi glider box, by source (2025–present)", fontsize=11)
ax.set_xlabel("datasets"); ax.spines[["top", "right"]].set_visible(False)
plt.tight_layout(); plt.show()
No description has been provided for this image

B4 · Find the glider — the way you'd actually find it¶

We don't know any glider IDs. We know a box and a time. That is enough: search the IOOS Glider Data Assembly Center inside the Hawaiʻi box and read what comes back.

In [8]:
gliders = list(catalog.search(bbox=HAWAII, collections=["IOOS"], datetime=WHEN).items())
print(f"{len(gliders)} glider deployments in the box since 2025\n")

for g in gliders:
    p = g.properties
    print(f"  {g.id:<24} {p.get('start_datetime','')[:10]} → {p.get('end_datetime','')[:10]}")
6 glider deployments in the box since 2025

  sg626-20250729T1452      2025-07-29 → 2025-10-29
  sg511-20260722T1831      2026-07-22 → 2026-09-14
  sg511-20250729T1532      2025-07-29 → 2025-09-09
  sg148-20260722T1939      2026-07-22 → 2026-09-14
  ng457-20260713T0000      2026-07-23 → 2026-09-14
  ng344-20260713T0000      2026-07-23 → 2026-09-14

We'll use sg626-20250729T1452 — a completed three-month Seaglider deployment, so the numbers in this notebook stay reproducible. (The 2026 deployments are still flying; their data grows every day. Swapping one in is the exercise at the end.)

Every item carries its variable list, so you can check it has what you need before downloading.

In [9]:
GLIDER = "sg626-20250729T1452"
g = catalog.get_collection("IOOS").get_item(GLIDER)

p = g.properties
print(p.get("title") or g.id)
print("period :", p["start_datetime"][:10], "→", p["end_datetime"][:10])

v = p["aquaview:variables"]
print(f"\n{len(v)} variables. The ones we want:")
for want in ["time", "latitude", "longitude", "depth", "temperature", "salinity"]:
    print(f"   {want:<12} {'✓' if want in v else '✗'}")

print("\nOther instruments on board:",
      ", ".join(x for x in v if x.startswith(("sbe43", "aa5013", "bb2f"))))
print("\nTabular assets:", ", ".join(k for k in ("csv", "nc", "json") if k in g.assets))
sg626-20250729T1452
period : 2025-07-29 → 2025-10-29

60 variables. The ones we want:
   time         ✓
   latitude     ✓
   longitude    ✓
   depth        ✓
   temperature  ✓
   salinity     ✓

Other instruments on board: aa5013_o2, bb2f_470, bb2f_695, bb2f_700, sbe43_o2

Tabular assets: csv, nc, json

B5 · Load the track¶

The csv asset points at the IOOS Glider DAC's own ERDDAP. AQUAVIEW standardises discovery and metadata; it does not hide the source — so we append ERDDAP's native tabledap query to pick columns and a time window, and only that slice is transferred.

Always constrain the time window. Without it you would ask for the entire three-month deployment.

In [10]:
T0, T1 = "2025-08-01T00:00:00Z", "2025-08-15T00:00:00Z"

url = (g.assets["csv"].href +
       "?time,latitude,longitude,depth,temperature,salinity"
       f"&time>={T0}&time<={T1}")

track = pd.read_csv(url, skiprows=[1], parse_dates=["time"])   # row 2 is ERDDAP's units row
track = track.dropna(subset=["latitude", "longitude", "temperature"])

print(f"{len(track):,} measurements over {(track.time.max()-track.time.min()).days} days")
print(f"  latitude   {track.latitude.min():.3f} → {track.latitude.max():.3f}")
print(f"  longitude  {track.longitude.min():.3f} → {track.longitude.max():.3f}")
print(f"  depth      {track.depth.min():.1f} → {track.depth.max():.1f} m")
print(f"  temperature {track.temperature.min():.2f} → {track.temperature.max():.2f} °C")
track.head()
84,176 measurements over 13 days
  latitude   18.827 → 20.994
  longitude  -158.135 → -156.176
  depth      -0.2 → 907.6 m
  temperature 4.61 → 29.38 °C
Out[10]:
time latitude longitude depth temperature salinity
0 2025-08-01 04:35:21+00:00 20.9936 -158.1345 0.60 27.390 NaN
1 2025-08-01 04:35:21+00:00 20.9936 -158.1345 0.76 27.373 NaN
2 2025-08-01 04:35:21+00:00 20.9936 -158.1345 0.72 27.368 NaN
3 2025-08-01 04:35:21+00:00 20.9936 -158.1345 0.67 27.358 NaN
5 2025-08-01 04:35:21+00:00 20.9936 -158.1345 1.01 27.393 NaN
In [11]:
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11, 4.2),
                               gridspec_kw={"width_ratios": [1, 1.35]})

sc = ax1.scatter(track.longitude, track.latitude, c=track.time.astype("int64"),
                 cmap="viridis", s=2, linewidths=0)
ax1.plot(track.longitude.iloc[0], track.latitude.iloc[0], "o", ms=7,
         mfc="none", mec="#b5642f", mew=1.6)
ax1.annotate("start", (track.longitude.iloc[0], track.latitude.iloc[0]),
             textcoords="offset points", xytext=(8, 2), fontsize=9, color="#b5642f")
ax1.set_xlabel("longitude"); ax1.set_ylabel("latitude")
ax1.set_title(f"{GLIDER} — track, 1–15 Aug 2025", fontsize=11)
ax1.set_aspect("equal", adjustable="datalim")
ax1.spines[["top", "right"]].set_visible(False)

s2 = ax2.scatter(track.time, track.depth, c=track.temperature,
                 cmap="RdYlBu_r", s=2, linewidths=0)
ax2.invert_yaxis()
ax2.set_ylabel("depth (m)"); ax2.set_title("Temperature section", fontsize=11)
ax2.spines[["top", "right"]].set_visible(False)
fig.colorbar(s2, ax=ax2, label="°C", pad=0.02)
for lbl in ax2.get_xticklabels():
    lbl.set_rotation(30); lbl.set_ha("right")

plt.tight_layout(); plt.show()
No description has been provided for this image

A two-week southbound transect from off Oʻahu down past Maui, diving to ~900 m. The section shows the warm mixed layer sitting on a sharp thermocline — exactly the structure that decides how sound travels, which is why it matters for passive acoustics.

B6 · Now check the satellite against it¶

This is the payoff, and it is the same comparison the CoastWatch glider tutorial makes — except we found both datasets through one catalog.

Step 1 — the satellite product. We know from B3 that SST lives in COASTWATCH. Find the blended analysis by searching inside our box rather than guessing an ID.

In [12]:
sst_items = catalog.search(bbox=HAWAII, collections=["COASTWATCH"], datetime=WHEN)
for it in sst_items.items():
    t = it.properties.get("title") or ""
    if "Blended" in t and "Sea-Surface Temperature" in t:
        print(f"{it.id:<26} {t[:64]}")
noaacwBLENDEDsstDaily      Sea-Surface Temperature, NOAA Geo-polar Blended Analysis Night O
noaacwBLENDEDsstDNDaily    Sea-Surface Temperature, NOAA Geo-polar Blended Analysis Day+Nig
noaacwBLENDEDsstDLDaily    Sea-Surface Temperature, NOAA Geo-polar Blended Analysis Diurnal
noaacwBLENDEDCsstDaily     Sea-Surface Temperature, NOAA Geo-polar Blended Analysis Night O

noaacwBLENDEDsstDaily is NOAA's Geo-polar Blended daily SST — geostationary plus polar-orbiting satellites merged onto a 5 km grid, 2002 to present. It is the night-only product, which is the right choice here: daytime skin warming makes the satellite look warmer than the top ten metres a glider actually samples, and night values compare far more fairly. (...DNDaily is the day+night version if you want to see that effect for yourself.)

Step 2 — open it lazily. The nc asset is a materialising ERDDAP endpoint: requesting it with no constraints asks for the whole 24-year record and ERDDAP refuses with 413 Payload Too Large. Strip the .nc extension and you get the griddap OPeNDAP endpoint, which xarray opens lazily — metadata only, nothing transferred until you subset.

In [13]:
import xarray as xr

item    = catalog.get_collection("COASTWATCH").get_item("noaacwBLENDEDsstDaily")
opendap = item.assets["nc"].href.removesuffix(".nc")     # ← the one character that matters

ds = xr.open_dataset(opendap)                            # lazy: nothing downloaded yet
print(opendap)
print("  grid :", dict(ds.sizes))
print("  units:", ds["analysed_sst"].attrs.get("units"))
https://coastwatch.noaa.gov/erddap/griddap/noaacwBLENDEDsstDaily
  grid : {'time': 8721, 'latitude': 3600, 'longitude': 7200}
  units: kelvin

Step 3 — subset to the track, then load. We take the glider's own bounding box plus a small pad, and only its two-week window. That turns a 24-year global grid into a few hundred kilobytes.

In [14]:
PAD = 0.25

sst = ds["analysed_sst"].sel(
    latitude =slice(track.latitude.min()  - PAD, track.latitude.max()  + PAD),
    longitude=slice(track.longitude.min() - PAD, track.longitude.max() + PAD),
    time     =slice(T0[:10], T1[:10]),
).load()                                     # ← the only network transfer

sst = sst - 273.15                           # the grid is in kelvin
print("subset:", dict(sst.sizes), "→", f"{sst.nbytes/1e3:.0f} kB")
subset: {'time': 15, 'latitude': 53, 'longitude': 49} → 156 kB

Step 4 — put them on the same footing. The glider samples continuously down to 900 m; the satellite sees one skin value per day. So we take the glider's shallowest 10 m and average it per day — the same choice the CoastWatch tutorial makes.

Watch the timezone. ERDDAP returns timezone-aware UTC; the grid's time axis is naive. Comparing them directly raises TypeError: Cannot compare dtypes datetime64[ns] and datetime64[us, UTC].

In [15]:
surface = track[track.depth <= 10].copy()
surface["time"] = surface["time"].dt.tz_localize(None)      # both naive UTC now

daily = (surface.set_index("time").resample("1D")
         .agg(lat=("latitude", "mean"), lon=("longitude", "mean"),
              glider_sst=("temperature", "mean"), n=("temperature", "size"))
         .dropna())

print(f"{len(surface):,} shallow measurements → {len(daily)} daily surface values")
daily.head()
1,942 shallow measurements → 14 daily surface values
Out[15]:
lat lon glider_sst n
time
2025-08-01 20.953288 -158.091301 27.336314 140
2025-08-02 20.847412 -157.983178 27.365098 143
2025-08-03 20.722256 -157.874002 27.492767 133
2025-08-04 20.592437 -157.744225 27.455970 134
2025-08-05 20.437748 -157.611933 27.348071 141

Step 5 — match each day to the satellite. For every glider day, take the satellite pixels within ±0.1° of where the glider actually was and average them. This is the match-up.

In [16]:
import numpy as np

def satellite_at(t, lat, lon, halfwidth=0.1):
    box = (sst.sel(time=t, method="nearest")
              .sel(latitude =slice(lat - halfwidth, lat + halfwidth),
                   longitude=slice(lon - halfwidth, lon + halfwidth)))
    return float(box.mean(skipna=True))

daily["sat_sst"] = [satellite_at(t, r.lat, r.lon) for t, r in daily.iterrows()]
daily["diff"]    = daily.glider_sst - daily.sat_sst

bias = daily["diff"].mean()
rmse = np.sqrt((daily["diff"] ** 2).mean())
corr = daily.glider_sst.corr(daily.sat_sst)

print(f"  mean bias    {bias:+.3f} °C   (glider minus satellite)")
print(f"  RMSE         {rmse:.3f} °C")
print(f"  correlation  {corr:.3f}")
print(f"  n            {len(daily)} days")
daily[["lat", "lon", "glider_sst", "sat_sst", "diff"]].round(3)
  mean bias    -0.039 °C   (glider minus satellite)
  RMSE         0.111 °C
  correlation  0.941
  n            14 days
Out[16]:
lat lon glider_sst sat_sst diff
time
2025-08-01 20.953 -158.091 27.336 27.358 -0.022
2025-08-02 20.847 -157.983 27.365 27.303 0.063
2025-08-03 20.722 -157.874 27.493 27.484 0.008
2025-08-04 20.592 -157.744 27.456 27.386 0.070
2025-08-05 20.438 -157.612 27.348 27.325 0.023
2025-08-06 20.282 -157.496 26.993 26.970 0.023
2025-08-07 20.137 -157.389 26.857 27.008 -0.151
2025-08-08 19.982 -157.263 26.824 27.001 -0.176
2025-08-09 19.797 -157.149 26.938 27.033 -0.094
2025-08-10 19.608 -157.021 27.055 26.966 0.089
2025-08-11 19.431 -156.870 27.363 27.280 0.083
2025-08-12 19.220 -156.688 27.198 27.260 -0.062
2025-08-13 19.042 -156.463 26.727 26.960 -0.233
2025-08-14 18.889 -156.238 26.840 27.003 -0.163
In [17]:
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11, 3.9),
                               gridspec_kw={"width_ratios": [1.5, 1]})

ax1.plot(daily.index, daily.glider_sst, "o-", color="#0b7d84", lw=1.6, ms=5,
         label=f"Glider {GLIDER.split('-')[0]} (≤10 m)")
ax1.plot(daily.index, daily.sat_sst, "s--", color="#b5642f", lw=1.6, ms=4.5,
         label="Satellite (Geo-polar Blended)")
ax1.set_ylabel("sea-surface temperature (°C)")
ax1.set_title("Glider vs satellite along the track", fontsize=11)
ax1.legend(frameon=False, fontsize=9)
ax1.spines[["top", "right"]].set_visible(False)
for lbl in ax1.get_xticklabels():
    lbl.set_rotation(30); lbl.set_ha("right")

lo = min(daily.glider_sst.min(), daily.sat_sst.min()) - 0.15
hi = max(daily.glider_sst.max(), daily.sat_sst.max()) + 0.15
ax2.plot([lo, hi], [lo, hi], "-", color="#c9c2b4", lw=1)
ax2.scatter(daily.sat_sst, daily.glider_sst, c="#0b7d84", s=34, zorder=3)
ax2.set_xlim(lo, hi); ax2.set_ylim(lo, hi); ax2.set_aspect("equal")
ax2.set_xlabel("satellite (°C)"); ax2.set_ylabel("glider (°C)")
ax2.set_title(f"bias {bias:+.2f} · RMSE {rmse:.2f} · r {corr:.2f}", fontsize=10)
ax2.spines[["top", "right"]].set_visible(False)

plt.tight_layout(); plt.show()
No description has been provided for this image

That is a real validation result, reproduced from a standing start in one notebook: the blended satellite product tracks this glider to about a tenth of a degree, with essentially no bias, over two weeks and 250 km of ocean. Where the two part company — around 7–8 and 13–14 August — the glider is cooler than the satellite, which is what you would expect when the wind mixes the surface and the skin value the satellite sees stops representing the top ten metres.

B7 · The layers a single-server tutorial can't add¶

Everything so far you could have done against one ERDDAP server. This is the part you cannot: the same box and window, asked of sources that have nothing to do with each other.

In [18]:
PAD = 0.25
track_box = [float(track.longitude.min() - PAD), float(track.latitude.min() - PAD),
             float(track.longitude.max() + PAD), float(track.latitude.max() + PAD)]
print("track box:", [round(v, 2) for v in track_box], "\n")

context = {
    "COASTWATCH":         "satellite SST / ocean colour",
    "PacIOOS":            "wave + circulation models",
    "MARINECADASTRE_AIS": "vessel traffic (noise context for PAM)",
    "OBIS":               "species occurrence datasets",
    "CERULEAN":           "SAR oil-slick detections",
    "IOOS":               "other gliders in the same water",
}
for cid, what in context.items():
    n = catalog.search(bbox=track_box, collections=[cid], datetime=WHEN).matched()
    print(f"  {n:>5}  {cid:<20} {what}")
track box: [-158.38, 18.58, -155.93, 21.24] 

     89  COASTWATCH           satellite SST / ocean colour
     49  PacIOOS              wave + circulation models
     26  MARINECADASTRE_AIS   vessel traffic (noise context for PAM)
     39  OBIS                 species occurrence datasets
    235  CERULEAN             SAR oil-slick detections
      6  IOOS                 other gliders in the same water

One box, six providers, one loop. For a PAM project the AIS row is the interesting one — vessel traffic is the dominant source of low-frequency noise, and it is discoverable here next to the oceanography rather than on a different agency's website.

Honest caveat, worth knowing on the day. The AIS records are in the catalog, but MarineCadastre publishes the underlying daily files about a year in arrears, and the 2025/2026 asset links in the catalog do not resolve yet. Treat AIS here as discovery — the catalog tells you what exists and when — and fetch 2023/2024 files directly from MarineCadastre if you need vessel data this week.

In [19]:
ais_hits = catalog.search(bbox=track_box, collections=["MARINECADASTRE_AIS"],
                          datetime="2025-08-01T00:00:00Z/2025-08-31T00:00:00Z").items()

for ais in ais_hits:
    print(f"{ais.properties.get('title'):<40} {len(ais.assets):>3} assets")

# the monthly record carries one asset per day of the month
monthly = next(a for a in
               catalog.search(bbox=track_box, collections=["MARINECADASTRE_AIS"],
                              datetime="2025-08-01T00:00:00Z/2025-08-31T00:00:00Z").items()
               if a.assets)
print("\n", monthly.id, "→", ", ".join(sorted(monthly.assets)[:5]), "…")
print(" example href:", monthly.assets["day_04"].href)
AIS Vessel Transit Counts 2025             0 assets
AIS Vessel Tracks 2025                     0 assets
AIS Vessel Traffic -- August 2025         32 assets

 marinecadastre_ais_2025_08 → data_browse, day_01, day_02, day_03, day_04 …
 example href: https://coast.noaa.gov/htdata/CMSP/AISDataHandler/2025/ais-2025-08-04.csv.zst

And the wave model, while we're here¶

The Hawaiʻi WaveWatch III run is in PacIOOS. Two things to get right, and they are the two mistakes everyone makes with ERDDAP:

  1. Pick the right longitude convention. PacIOOS publishes each model twice. ww3_hawaii runs 0–360° (199→206 here); ww3_hawaii_lon180 runs −180–180° and matches our box exactly.
  2. Give every axis a constraint. Thgt is [time][depth][latitude][longitude] — four axes. Supply three and ERDDAP silently reads your latitude range as depth and returns an error image.
In [20]:
import requests
from IPython.display import Image, display

wave = catalog.get_collection("PacIOOS").get_item("ww3_hawaii_lon180")
print(wave.properties["title"])
print("variables:", wave.properties.get("aquaview:variables"))

graph = wave.assets["png"].href + (                     # Thgt[time][depth][lat][lon]
    "?Thgt[(last)][(0.0)][(18.0):(23.0)][(-161.0):(-154.0)]"
    "&.draw=surface&.vars=longitude|latitude|Thgt&.land=under"
)
r = requests.get(graph, timeout=60)

# ERDDAP reports failure in the *body* with HTTP 200 — check the type, not the status code.
if "image" in r.headers.get("content-type", ""):
    display(Image(data=r.content))
else:
    print("ERDDAP said:", r.text[:300])
WaveWatch III (WW3) Hawaii Regional Wave Model, Lon+/-180
variables: ['Tdir', 'Tper', 'Thgt', 'sdir', 'sper', 'shgt', 'wdir', 'wper', 'whgt']
No description has been provided for this image

B8 · On the map, and by just asking¶

Prefer clicking? The same box and sources on the live Explore map. The screenshot links to the real, preconfigured view — note that Explore counts what is in the current viewport, so its number differs from .matched() above, which counts the whole box.

AQUAVIEW Explore — Hawaiʻi box, 5 sources, hex-binned

Prefer asking? AQUAVIEW's MCP server is live at https://mcp.aquaview.org/mcp. Point any MCP client at it — in Claude Desktop, add:

{ "mcpServers": { "aquaview": { "url": "https://mcp.aquaview.org/mcp" } } }

then ask "what SST and vessel-traffic data is near 20°N 157°W in August 2025?" and you get the same STAC items you just found by hand. Same catalog, third door.


Your turn¶

You now have the whole loop: meet the catalog → search a place and time → narrow → count → find an instrument → load it → validate against it.

Change one thing and re-run:

  1. Another glider. Swap GLIDER for sg511-20260722T1831 — it is flying right now. You'll need to move T0/T1 into its window (B4 prints the dates). Does the satellite agree as well?
  2. Another variable. salinity is already in the track DataFrame. COASTWATCH also carries ocean colour — try matching chlorophyll instead of SST.
  3. Another place. Change HAWAII to a box over your own study area and re-run B1–B4. Which sources show up that you didn't expect?
  4. Tighten the match-up. We used a ±0.1° box and a daily mean. Does the RMSE change if you use the single nearest pixel instead?

For a hackathon project, the sources to reach for:

Question Source
Ocean state along the track COASTWATCH (SST, colour), PacIOOS (waves, currents)
What vessel traffic is nearby? MARINECADASTRE_AIS (discovery; files lag ~1 year)
Which species are recorded here? OBIS
Sound-speed / CTD context the glider's own temperature + salinity, WOD, GADR
Other gliders in the same water IOOS

Free account, STAC access included → aquaview.org