Seantral: how it works inside, and the things that broke
· 8 min · marine, forecast, edge, testing
Marine forecasts turned into an honest verdict. The map was the easy part. The hard part was not lying when the data goes missing.
It started from a dull question that came up every time before going out: what’s it actually like today, at this exact spot? Weather sites give you numbers. Wind 14 knots, wave 0.8 metres. But turning those numbers into “worth it or not” was mental arithmetic, done every time and every time slightly differently. The idea: that translation step can be written once and kept honest.
Where the numbers come from
The data is Open-Meteo (Weather + Marine, CC-BY 4.0). The thing few people do, and the thing that matters here: no single model. Four weather and three wave models are blended per variable, and instead of a flat average the distribution is kept: mean, spread, percentiles. The spread between models isn’t thrown away: it becomes the confidence shown on screen. When the models agree, the verdict is sharp. When they fight, the app says so.
An hourly cron refreshes the cache for all the curated spots, in one batch. The client reads from the cache, it doesn’t call the providers from the browser. Sounds like a detail but it changes two things: the map is fast, and the providers don’t know who’s looking at what.
A couple of precautions that look like paranoia until you need them. When the cache expires and needs refreshing, a single requester does the fetch while the others wait on its result, instead of all firing at once and hammering the provider (anti-stampede). And if the fetch fails, the last cache is served flagged as stale, with the notice: never a silent gap. The map downloads a compact per-spot summary; the detail loads only when you open the card.
The verdict, and why it leans pessimistic
The engine takes wave, wind, period, shelter and time-of-day and folds them into a 0–10 score with a word on top (calm, choppy, rough). Each activity weighs the same data differently: swimmers fear the wave, boaters watch the wind, anglers tolerate some chop. No science was invented here: they’re “ideal/worst” thresholds per factor, hand-tuned and still under field validation, that push middling scores down rather than rounding them up. Better to say “choppy” and find flat water than the other way round.
The rule never broken here: the score is a comfort read, not a safety assessment. For deciding whether to go out, the official notices are what count, full stop. The sentence appears in three separate places in the app because it’s the one that must leave no ambiguity.
Make the data vanish, not the signal
Behind a single number there’s more stuff than whoever’s looking wants to see: seven models blended into distributions, the per-factor thresholds, the capped carried-forward wind, the missing-data handling. When you tap a spot, none of it reaches you. What reaches you is a word (calm, choppy, rough), a 0–10 score and the best window over the next hours. The hard part of the project wasn’t handling the data. It was hiding it without throwing the signal away.
The map is the whole app. Colored badges per spot, readable at a glance, with no form to fill and no setting to pick. And the activity lens: swimming, boating, fishing re-fish the same data with one tap and change the reading, without adding a gram of complexity on screen. Underneath, all the work; on top, one switch.
The rest is there for whoever goes looking. The map shows a compact summary; the factor breakdown and the hourly trend load only when you open the spot, so the first view stays light and fast. And honesty doesn’t have to shout: the amber missing-data flag and the confidence coming from model spread sit there, discreet, for whoever wants them, and they don’t jump at you on the first look.
The AI features live by this same rule. The assistant answers “where to swim” from the same engine as the map, so chat and map never contradict each other on a spot. The on-site report is left in three taps (as forecast, better, worse) and closes the loop toward the future scorecard. And the raw feedback gets classified by the AI into change requests, so the reports don’t pile up unread. The point is always the same: a lot of work underneath, one single thing to look at on top.
The evening of “4.7” on a flat sea
This one cost an evening. The app showed a mediocre score for a spot while outside it was a mirror. The evening went on hunting a bug in the engine, and the engine was fine. The problem was upstream: Open-Meteo sometimes has nightly holes on the 3–6 hours, and the wind was missing for that window. The code carried the last known wind forward, sensible enough, but with no cap, and without saying so.
The fix made the hole honest instead of hiding it: carried wind has a six-hour cap from the data’s origin; beyond that it explicitly degrades to “wind missing”, the score is capped around 5, and an amber notice shows up. No factor without signal gets penalised silently: if a value is missing, you can see it’s missing. The number always travels with its word, even when the word is “don’t trust this too much”.
One engine, two runtimes, zero “it’s different on this device”
The verdict computation is needed in two places: in the browser (to answer instantly when you tap a spot) and on the edge (when the AI assistant answers “where to swim”). The temptation is to write the same logic twice. That’s exactly how, three patches later, the map and the AI hand you two different numbers for the same spot, and nothing crashes: they just drift. A nightmare to find.
The answer was a single pure module, shared between client and edge, locked with golden vectors: a set of cases with frozen input and expected output, run in CI on both runtimes: Vitest on the browser side, deno test on the edge. If a change makes the two diverge, the build fails before anyone notices by hand. It’s the net that made it possible to touch the engine without fear. And in fact, when the AI was realigned to the spot list, the bug was right there: the AI used the nowcast frozen at fetch time instead of projecting to the current hour like the map does. Same engine, different reference hour, different ranking.
The report from whoever’s actually there
A forecast is a forecast. What it lacks is someone in the water saying “yes, it’s like the app says” or “no, it’s worse”. Hence a quick report: how it is versus the forecast: same, better, worse. With one condition: to leave it you have to actually be there. The client computes the distance from the spot (Haversine, threshold around 3 km) and with no location, or too far, the report doesn’t go through. What gets stored is the declared distance, not the coordinates. It’s a soft check, not a burglar alarm: it keeps reports anchored to reality, it doesn’t track anyone.
Where reports diverge sharply from the forecast, that’s a bridge to the part that matters most long-term: the per-spot scorecard. Forecast versus observed, and over time the formulas get tuned on real data instead of by eye. For now the numbers are few and it’s all under validation. It’s worth saying plainly, because the worst thing would be passing an experiment off as science.
Where it’s heading
The direction is two steps. First collect real signal: the on-site reports, the divergences between models, the curated webcams. Then use it to recalibrate the formulas on data instead of by eye: a per-spot scorecard comparing forecast against observed over time. On top there are two experiments on the list: a multi-factor bathing index (water, air, UV, wind) and a “dirty sea” read from public water-quality sources, with a turbidity flag after it’s just rained.
Further out, the most speculative part: a measurement layer on site, not just forecast. Existing tide-gauge networks, in-situ data, and (as a demonstrator, not a product) a home-made buoy with a microcontroller, motion sensors, GPS and a small solar panel. The value lies in comparing what the model says with what the sea actually does, more than in the hardware. It’s an exploration, and it’s marked as such.
What it taught
That the hard part of a forecast app is the behaviour when the data is missing. An engine that always returns a full number is simpler to write and easier to distrust. All the real work was building the honest ways of saying “unknown”: the carry-forward cap, the visible missing factor, the confidence from model spread. And a single engine, tested across two runtimes, because trust goes fast the day two screens of the same app disagree.