PAM–Glider Rodeo Hackweek · Tue Sep 15, 2–3 PM Central · 60 min · live + recorded

Discovering ocean data with AQUAVIEW

A hands-on onboarding tutorial. First we meet the tool — one common catalog over 600,000+ ocean datasets from ~100 sources, reachable three ways. Then we put it to work on a real question: is the satellite right, along a Hawaiʻi glider track?

STAC API confirmed live · Python + R …/stac · unauthenticated · pystac-client + rstac ready Storyboard for review — not the final notebook
608,400
datasets in one catalog, worldwide
90
ocean & atmosphere sources, one API
4,228
in the Hawaiʻi box −161,18,−154,23, 2025–now
R + Python
same workflow, both languages
Part A

Meet AQUAVIEW

Before the glider example, orient the room: what AQUAVIEW is, how much is in it, how you find things, and the three doors into the same catalog. ~14 minutes.

Meet the tool. This is aquaview.org — the front door: one platform, one catalog, with an Explore map, an API, and AI search over it. Today we drive it from code.

aquaview.org homepage — one data plane from ingest to access, with the Explore map of global data density

Under the hood — one catalog, three ways in. ~100 sources feeding in; today we learn the Code door, but the same catalog is behind all three.

AQUAVIEWone common catalog · 600,000+ datasets · ~100 sourcesExplorewebsite · point & clickCodePython · R (STAC API)AIMCP · plain Englishthe same catalog underneathNOAAIOOSArgoOBISmodelsAISbuoyssatellite
Five words we’ll use
AQUAVIEW
   └─ Source / Collection    e.g. "PacIOOS"     a data provider or product family
        └─ Dataset / Item    e.g. "ww3_hawaii"  one dataset record, with location + time
             └─ Assets       nc · csv · png     the actual downloadable files
A1

Connect — the whole catalog is one object

AQUAVIEW speaks the STAC standard, so the standard client just works — one line in the language you already use, no sign-in for discovery.

Pythonmeet.py
from pystac_client import Client

STAC = "https://service.aquaview.org/stac"
cat = Client.open(STAC)
print("Connected to AQUAVIEW ✓")
Rmeet.Rmd
library(rstac)

stac_url <- "https://service.aquaview.org/stac"
cat <- stac(stac_url)
cat("Connected to AQUAVIEW ✓\n")
A2

One catalog over everything

AQUAVIEW federates the whole ocean-data landscape into one searchable place. 608,400 datasets, ~100 sources. Instead of tracking down data across many separate servers, clouds and portals, you discover and reach it in one place.

Output
608,400 datasets  ·  100 sources  ·  one catalog, one API

Ship & float profiles

WOD · Argo (GADR) · EN4 · CCHDO

Satellite — SST, colour, SAR

CoastWatch · GOES-R · Cerulean oil-slicks

Gliders & autonomous

IOOS Glider DAC · Spray · Voice of the Ocean

Buoys & sensors

NDBC · IOOS Sensors · CDIP waves

Ocean & weather models

HYCOM · RTOFS · Copernicus GLORYS

Biodiversity & human activity

OBIS species · MarineCadastre AIS

A3

Find a source without memorising IDs

The transfer skill: you won’t know collection IDs cold. Each source carries a title, description and keywords — already fetched once in A2 — so we filter locally: no extra calls, and you can see exactly what matched. Read the result carefully: it searches how a source describes itself, so find_sources("wave") misses PacIOOS. Rule: a provider you half-remember → find_sources(); datasets in your study area → a bounding-box search (Part B); browse → Explore.

Pythonmeet.py
# local filter over id+title+description+keywords
HAY = {c.id: haystack(c) for c in collections}
def find_sources(term, n=6):
    return [cid for cid, h in HAY.items()
            if term.lower() in h][:n]
Rmeet.Rmd
# rstac: free-text the fetched list over
# id + title + description + keywords
find_sources("glider")
Output
glider      → CALOOS, CCHDO, CENCOOS, IOOS, MARACOOS, NANOOS
wave        → CDIP, CENCOOS, GFS, NDBC
chlorophyll → NEFSC          # matched via keywords, no ID known
vessel      → CERULEAN, MARINECADASTRE_AIS, …
A4

Three ways in — same catalog, your choice

Whichever door you come in, a result is the same STAC item with the same schema — so the data looks identical downstream.

No code

Explore map

Click-and-draw discovery at aquaview.org/explore — pan to an area, set a time slider, see what’s there.

for → browsing, demos, newcomers
Scriptable

STAC API

pystac-client / rstac against the live endpoint — reproducible, versionable, drops into a notebook.

for → analysis, the rest of today
Plain English

AI / MCP chat

Ask a question; the AI layer turns it into the same STAC query — no need to know a collection ID.

for → “what’s near my track?”
Part B

Worked example — is the satellite right, along a Hawaiʻi glider track?

Now put the tool to work. Your gliders record passive acoustics and CTD off Hawaiʻi; before you can interpret that you need the ocean around the track — and to know whether the satellite products you lean on agree with what the glider measured. Six moves: search a box → narrow → count → find the glider → load its track → match the satellite to it. ~38 minutes — everything typed live.

B1

Search one place, one window

The core move: a bounding box around the fleet and a date range. One call reaches across every source — no per-server hunting.

Pythonglider.py
HAWAII = [-161, 18, -154, 23]   # W,S,E,N
hits = cat.search(bbox=HAWAII,
                  datetime="2025-01-01/..")
print(hits.matched(), "datasets")
Rglider.Rmd
hits <- cat |> stac_search(
    bbox = c(-161, 18, -154, 23),
    datetime = "2025-01-01T00:00:00Z/.."   # R: full RFC3339
  ) |> get_request()
Output
4,228 datasets match  # 2025–present, across the Hawaiʻi box
B2

From a broad catch to what you actually want

A wide search returns everything intersecting your box and time — that’s discovery working, not failing. Here it’s dominated by Cerulean SAR oil-slicks. So you narrow with collections= — the real skill — and pull one dataset per relevant source. Same columns, wildly different sources. Mind the SST source: satellite SST lives in plain COASTWATCH; COASTWATCH_WC is the West-Coast node — over Hawaiʻi it carries HF-radar currents and ocean colour.

Output — broad catch (unfiltered)
source     id        title
CERULEAN   5214248   Oil slick 5214248
CERULEAN   5214247   Oil slick 5214247       # dense here — but not what a glider needs
CERULEAN   5214246   Oil slick 5214246
Output — narrowed to relevant sources
source              id                   title
COASTWATCH_WC       ucsdHfrH1_Lon0360    Currents, HF Radar, US Hawaiʻi State
PacIOOS             ww3_hawaii_lon180    WaveWatch III (WW3) Hawaiʻi Wave Model
OBIS                345c1b88-2bab…       DFO Maritimes Cetacean Sightings
MARINECADASTRE_AIS  mc_dataset_67919     Vessel Routing Measures
IOOS                sg626-20250729…      Seaglider sg626 track
›The narrowed table even surfaces a real Seaglider track — so the example touches actual glider data, not just context layers.
B3

What’s here, by source?

Before downloading anything, ask how much of each kind exists in your box — one cheap count per source, no data moved. This is the bar chart in the live notebook.

Output · matched() counts, Hawaiʻi box 2025–now
Ocean colour / HF radar (COASTWATCH_WC)  ████████████  293
Satellite SST (COASTWATCH)               ████          91
Hawaiʻi waves & models (PacIOOS)        ███           61
Species-occurrence context (OBIS)        ██            42
Vessel-traffic context (AIS)             █             27
Glider Data Assembly Center (IOOS)       █              6
›Counts are honest: ocean colour / HF-radar is richest in-box and satellite SST (plain COASTWATCH) has 91; vessel-traffic and glider layers are sparser (coverage/indexing, not absence). One matched() per collection — totals without fetching a file.
B4

Find the glider — the way you’d actually find it

We don’t know any glider IDs. We know a box and a time — that’s enough. Search the IOOS Glider DAC inside the Hawaiʻi box and read what comes back. Every item carries its variable list, so you check it has what you need before downloading.

Pythonglider.py
gliders = list(cat.search(bbox=HAWAII,
    collections=["IOOS"], datetime=WHEN).items())
for g in gliders:
    p = g.properties
    print(g.id, p["start_datetime"][:10],
          "→", p["end_datetime"][:10])
Rglider.Rmd
gl <- cat |> stac_search(collections = "IOOS",
        bbox = hawaii, datetime = when) |> get_request()
for (f in gl$features)
  cat(f$id, f$properties$start_datetime, "\n")
Output
6 glider deployments in the box since 2025
  sg626-20250729T1452   2025-07-29 → 2025-10-29   # completed 3-month Seaglider — our pick (reproducible)
  sg511-20260722T1831   2026-07-22 → 2026-09-14   # flying right now — the “your turn” swap
  sg511-20250729T1532 · sg148-20260722T1939 · ng457-20260713T0000 · ng344-20260713T0000

sg626: 60 variables — time ✓ latitude ✓ longitude ✓ depth ✓ temperature ✓ salinity ✓
       also aboard: sbe43_o2, aa5013_o2, bb2f_470/695/700   ·   tabular assets: csv, nc, json
B5

Load the track

The csv asset points at the Glider DAC’s own ERDDAP. AQUAVIEW standardises discovery and metadata; it doesn’t hide the source — so we append ERDDAP’s native tabledap query to pick columns and a time window, and only that slice is transferred. Always constrain the time window — without it you’d ask for the whole three-month deployment.

Pythonglider.py
T0, T1 = "2025-08-01T00:00:00Z", "2025-08-15T00:00:00Z"
url = (g.assets["csv"].href +
  "?time,latitude,longitude,depth,temperature,salinity"
  f"&time>={T0}&time<={T1}")
track = pd.read_csv(url, skiprows=[1],       # row 2 = units
                    parse_dates=["time"])
Rglider.Rmd
url <- paste0(g$assets$csv$href,
  "?time,latitude,longitude,depth,temperature,salinity",
  "&time>=", T0, "&time<=", T1)
track <- readr::read_csv(url, skip = 2, col_names = ...)
Output
84,176 measurements over 13 days
  latitude    18.827 → 20.994      longitude  −158.135 → −156.176
  depth        −0.2 → 907.6 m      temperature  4.61 → 29.38 °C
›A two-week southbound transect off Oʻahu past Maui, diving to ~900 m. The section shows the warm mixed layer on a sharp thermocline — the structure that decides how sound travels, which is why it matters for PAM.
B6

Now check the satellite against it — the payoff

The same comparison the CoastWatch glider tutorial makes — except both datasets came through one catalog. Five steps: find the blended SST by searching the box in COASTWATCH (noaacwBLENDEDsstDaily, night-only — fair against a glider’s top 10 m); open it lazily via OPeNDAP (strip .nc — the one character that matters; unconstrained it 413s); subset to the track box + window; daily-mean the glider’s shallowest 10 m; match ±0.1° pixels per day.

Pythonglider.py
item = cat.get_collection("COASTWATCH").get_item("noaacwBLENDEDsstDaily")
ds  = xr.open_dataset(item.assets["nc"].href
                      .removesuffix(".nc"))     # OPeNDAP, lazy
sst = ds["analysed_sst"].sel(latitude=…, longitude=…,
      time=slice(T0[:10], T1[:10])).load() - 273.15
daily = track[track.depth <= 10].resample("1D", on="time").mean()
daily["sat"] = [satellite_at(t, r.lat, r.lon)      # ±0.1° box
                for t, r in daily.iterrows()]
Rglider.Rmd
nc <- nc_open(sub("\\.nc$", "", item$assets$nc$href))
# ncdf4 reads by index: find the lat/lon/time index ranges,
# ncvar_get(start=, count=) only the track box + window,
# then the same daily ≤10 m mean and ±0.1° match-up.
Output
mean bias    −0.039 °C   (glider minus satellite)
RMSE          0.111 °C
correlation   0.941
n             14 days
Glider vs satellite SST along the sg626 track — time series and scatter; bias −0.04, RMSE 0.11, r 0.94
›A real validation result from a standing start. The blended product tracks this glider to about a tenth of a degree with essentially no bias, over two weeks and 250 km. Where they part company (7–8 and 13–14 Aug) the glider is cooler — wind mixes the surface and the skin value the satellite sees stops representing the top ten metres.
B7

The layers a single-server tutorial can’t add

Everything so far you could have done against one ERDDAP server. This you cannot: the same box and window, asked of sources that have nothing to do with each other — one loop.

Pythonglider.py
track_box = [lon.min()-PAD, lat.min()-PAD, lon.max()+PAD, lat.max()+PAD]
for cid in ["COASTWATCH","PacIOOS","MARINECADASTRE_AIS",
            "OBIS","CERULEAN","IOOS"]:
    n = cat.search(bbox=track_box, collections=[cid],
                   datetime=WHEN).matched()
Rglider.Rmd
for (cid in providers)
  n <- (cat |> stac_search(bbox = track_box,
         collections = cid, datetime = when, limit = 1) |>
         get_request())$numberMatched
Output · track box [−158.38, 18.58, −155.93, 21.24]
   89  COASTWATCH           satellite SST / ocean colour
   49  PacIOOS              wave + circulation models
   26  MARINECADASTRE_AIS   vessel traffic (noise context for PAM)
   39  OBIS                 species occurrence datasets
  235  CERULEAN             SAR oil-slick detections
    6  IOOS                 other gliders in the same water
›Honest caveat for the day: the AIS records are in the catalog (the Aug-2025 record carries 32 day-assets), but MarineCadastre publishes the daily files ~1 year in arrears and the 2025/26 links don’t resolve yet — treat AIS as discovery. Wave model: use ww3_hawaii_lon180 (−180…180 matches our box, not the 0–360 ww3_hawaii) and constrain all four axes Thgt[time][depth][lat][lon] — give ERDDAP three and it silently reads your latitude as depth.
B8

On the map — the same box, live

The bounding box is the search area. Left: the concept — your box and a glider track over the islands. Right: that exact view live on Explore (note it counts what’s in the viewport, so its number differs from matched(), which counts the whole box). Click to open the real map.

conceptYour search area — bbox + glider track
bbox −161,18,−154,23 glider track
B8

…and by just asking — the MCP server is live

AQUAVIEW’s MCP server is live at https://mcp.aquaview.org/mcp. Point any MCP client at it — in Claude Desktop, add it under mcpServers — then ask in plain English and you get the same STAC items you just found by hand. Same catalog, third door.

PythonClaude Desktop config
{ "mcpServers": { "aquaview": {
    "url": "https://mcp.aquaview.org/mcp" } } }
Ror the web chat
# literally, in the AQUAVIEW chat:
"What SST and vessel-traffic data
 is near 20°N 157°W
 in August 2025?"
Output
SST · HF-radar currents · AIS vessel traffic — described in plain language,
using the same items you searched by hand. Same catalog, third door.

Your turn — bring it back to your project

Whole loop: meet the catalog → search → narrow → count → find a glider → load it → validate — by code, by click, or by asking. Then get them to change one thing.

Try it — 60 seconds

1. Swap GLIDER for sg511-20260722T1831 — it’s flying right now (move T0/T1 into its window). Does the satellite agree as well? 2. Match chlorophyll instead of SST. 3. Change HAWAII to your own study area. 4. Tighten the match-up: nearest pixel instead of a ±0.1° box — does RMSE change?

Vessel context

AIS traffic

MarineCadastre AIS — vessel-traffic context for PAM. Discovery now; the underlying daily files lag ~1 year.

Who’s present

OBIS occurrences

Marine-species records — occurrence context for what your PAM detections might be.

Ocean state

SST · currents · waves

CoastWatch SST, HF-radar currents, PacIOOS waves; WOD/Argo for sound-speed.

The hour, minute by minute

A full hour (Tue Sep 15, 12:00–13:00 PT / 2–3 PM Central), so v2’s Part B fits comfortably: everything is typed live, the satellite-vs-glider validation gets twelve unhurried minutes, and there is a real hands-on block before Q&A.

0–3 min
Hook + the one picture → 0:03
Why a glider pilot needs the ocean around the track — and whether the satellite can be trusted. The one-catalog / three-doors diagram; the homepage as the front door.
3–9 min
Part A · Connect + breadth → 0:09
A1 connect (one host, defined once); A2 608k datasets, ~100 sources, the six kinds — one live search.
9–14 min
Part A · Find a source + three ways → 0:14
A3 local keyword filter and why it misses PacIOOS; A4 Explore live, then STAC + MCP as the other two doors.
14–21 min
Part B · Search → narrow → count → 0:21
B1 bbox (4,228); B2 broad catch → narrowed one-per-source table, and the COASTWATCH vs COASTWATCH_WC point; B3 the by-source bar chart.
21–29 min
Part B · Find the glider, load the track → 0:29
B4 list the 6 IOOS deployments in the box, pick sg626, check its variables; B5 tabledap CSV with a time window → track map + temperature section.
29–41 min
Part B · Validate vs satellite — the payoff → 0:41
B6 typed step by step: find the blended SST in COASTWATCH, open lazily via OPeNDAP (the .nc strip), subset, daily top-10 m, ±0.1° match-up → bias −0.04 · RMSE 0.11 · r 0.94 — and what the two divergences mean.
41–48 min
Part B · Layers a single server can’t add → 0:48
B7 six providers in the track box in one loop; the AIS honesty caveat; the wave model with the lon180 + four-axis gotchas.
48–52 min
Part B · Map + MCP, live → 0:52
B8 the deep-link Explore view (viewport count vs matched()); point an MCP client at mcp.aquaview.org/mcp and ask the question in plain English.
52–60 min
Your turn + Q&A → 1:00
Hands-on: swap in the glider that’s flying now (sg511), match chlorophyll, or change the box to your own study area. Then questions.

What the JupyterHub needs

Derived from the exact code above — the package list for Shannon to pre-install so every participant’s notebook runs on the day. Both languages confirmed working live against the catalog. Held until you approve this preview.

Python3.10+
  • pystac-client
  • pandas
  • xarray
  • netCDF4
  • requests
  • matplotlib
R4.x
  • rstac
  • readr
  • ncdf4
  • ggplot2
  • dplyr
  • rmarkdown

Before Sept 15

✓
Clean endpoint live — service.aquaview.org/stac
Both notebooks now use it; the run.app URL is gone from every file.
Done
✓
v2 notebooks adopted as canonical — Python & R (both IRkernel/ipynb), executed clean
Find glider → load track → validate vs satellite: bias −0.039 · RMSE 0.111 · r 0.941 · n 14 — reproduced in our environment.
Done · this storyboard matches them
✓
Hub packages confirmed & notebooks going into hackathon-shared-repo
Shannon: all R/Python packages already in the Hub mirror. Kapil: added updated reqs, adding the notebooks as “tutorials on aquaview”.
Done · via Slack
✓
Slot confirmed as a full hour — run-of-show re-timed, nothing cut
Everything typed live; validation gets 12 min; 8-min hands-on + Q&A block at the end.
Done
□
One end-to-end dry run before Tuesday
Type B4–B7 against the live catalog once at full speed — the OPeNDAP open and the tabledap pull are the two network beats to time.
You + Bibas