drift 0.14.1
-
Temporal QA across four watershed groups (#62).
dft_rast_break_class()run on the published seven-year IO LULC series ofbulk_co_ff04,necr_ch_ff04,lnth_ch_ff04andkotl_bt_ff04straight from stac-floodplains-bc: the share of 2017 -> 2023 change that is a switch sustained two years each side is 19.7-31.0%, flicker 39.6-48.5%, and 2017 is the odd endpoint more often than 2023 in every group. Results, what they do and do not support, and the two follow-up decisions are ininst/notes/temporal-qa-groups.md; evidence indata-raw/logs/break_class_groups/, every number emitted bydata-raw/break_class_groups.R. No API change.
drift 0.14.0
-
New:
dft_rast_break_class()scans the annual class series per pixel for a sustained switch versus flicker, and dates the switch (#9). This is the temporal, categorical-only leg of change QA, beside the spatial one (dft_transition_artifact(), #44) and the spectral one (dft_rast_break(), #30). For each pixel across a named list of classified years it reportsn_flips(0 stable, 1 a clean switch, 2 or more flicker),break_year(the first year of the new class), andn_before/n_after(years in the old and new class either side of it — confidence is their minimum). Its$rasteris the first-year-to-last-year transition layer in the samefrom * 1000 + toencoding asdft_rast_transition(), sodft_transition_vectors()anddft_transition_artifact()take it unchanged;$breaksholds the four evidence layers and$summarytabulates area by transition, status and break year. Nothing is thresholded: the caller composes, e.g.n_flips == 1 & pmin(n_before, n_after) >= 2. -
Why a mode filter cannot do this.
dft_rast_consensus()votes a pixel that genuinely switched mid-window back to its old class whenever the pre-change years outnumber the post-change ones, and cannot say when anything changed. The scan keeps the switch and dates it, and separates it from label noise that the two-epoch comparison reports as change. -
Measured on the whole BULK floodplain (the
co_ff04polygon of stac-floodplains-bc, 386.5 km², IO LULC 2017-2023 fetched throughdft_stac_fetch(tile_size = 20000): seven years in 23.6 min, a 16000 x 12000 grid at 10 m, 192M cells): the 2017 -> 2023 comparison reports 4,620 ha of change in 21,710 patches. Only 19.7% of that area is a switch sustained for at least two years on each side. 36.4% is a clean switch whose new or old class holds for a single year — 2017 or 2023 is itself the odd year out (922 ha broke in 2018, i.e. 2017 alone differs; 757 ha in 2023) — and 44.0% flickers: the label went back and forth and never settled. A further 3,187 ha flickers while reading stable on the endpoints, invisible to any two-epoch table. On the bundled Neexdzii Kwa tile (34 ha of change) the split is 32% sustained, 31% endpoint-only, 37% flicker. -
The temporal leg and the geometric leg are mostly independent, which is the point of having both. On BULK, patches carrying the #44 artifact signature (
sliver & (boundary | reciprocal), 15,047 patches, 826 ha) have an area-weighted clean-break share of 0.50 against 0.58 for the rest; 46% of them contain no clean-break cell at all, but 31% are entirely clean — one-pixel bands that moved once and stayed. Geometry ranks what to look at; the years say whether it settled. -
Scale numbers. One streamed
terra::app()pass to an LZW-compressed INT4S temp file and onecrosstab(); nothing full-grid enters R. BULK scan 90 s (87-101 s across runs); the whole pipeline (fetch from cache, classify, scan, vectors, artifact tags, per-patch zonal) ran in 6.2 min. Bundled tile: 0.15 s. Evidence:data-raw/logs/benchmark_break_class/, every number above derived from its CSVs by the committed script. -
Seven review rounds found ten defects the suite could not, all of the same shape: a terra internal contract assumed rather than measured.
terra::app()triesapply(chunk, 1, fun)— one R call per cell — beforefun(chunk), and afunthat tolerates a bare vector silently runs per cell (57x slower; the scan closure now refuses anything but a matrix, and the benchmark script reproduced the same bug in its own helper).INT2Soverflows at class code 33 (33 * 1000 + 33 > 32767, i.e. every ESA WorldCover code) and terra warns rather than errors, writingNA— output is INT4S and the warning is promoted to an abort.resample()aligns grids but does not reproject, so a different projected CRS was scanned by cell position — refused now.app()infers output shape from a test chunk ofmin(ncol, 13)cells and reads a 5-column return on a 5-column raster as transposed, scrambling cells across layers with no warning — such a stack is padded by one column for the scan. The overflow abort fires fromwriteValues()beforewriteStop(), so the partial output had to join the cleanup list.coltab<-on a stack strips layer 1 only, andlevels<-deep-copies before it strips, so the in-place formset.cats(NULL)is the one a caller-unmutated test can guard.resample()of a factor writes a RAT sidecar beside the intermediate. Andpaste0(character(0), ".aux.xml")is".aux.xml", which an unguarded cleanup would unlink in the caller’s working directory. Round 3’s deliverable was an enumeration of every terra call in the function with its assumption and how it was measured; it lives in the archived planning directory. -
Memory. On in-memory inputs, per-layer
deepcopy()and thenrast(list)each copied the whole series, andapp()left to its memory heuristic takes a 192M-cell grid in one or two chunks whose R-side matrices alone cost ~10 GB. In-memory inputs are now spilled to LZW temp files before stacking (a one-layer transient; the caller’s rasters keep their levels and colours — pinned), levels are stripped in place on the stack, andwopt$stepsbounds chunks to 2.5M cells. Measured on BULK with RSS sampled every 2 s: the seven fetched inputs set an 8.5-20 GB floor that belongs to the caller; from a 0.3 GB file-backed floor the scan peaked at ~10 GB with 9.6M-cell chunks and the crosstab was flat at 7 GB; on the shipped code with in-memory inputs the spill peaks at 20.9 GiB for a few seconds and the scan runs at 12-19 GiB — that spill sample is also the whole-pipeline peak (RSS in KiB / 1024², the same convention as #44’s 10.7 GB on a smaller grid from pre-clipped files). - The bundled example series now covers every IO LULC year 2017-2023 on the one grid (
data-raw/example_years_extend.R; a refetched 2017 reproduced the bundled file cell for cell), and the land-cover vignette gains a Temporal Evidence: Switch or Flicker? section as the third leg besidepatch_area_minand the geometric tags.
drift 0.13.0
-
New:
dft_transition_artifact()tags transition patches with geometric misregistration evidence, and drops nothing (#44). Year-to-year classified land cover carries artifacts that read as change but are misalignment: a class boundary that moved one pixel between epochs leaves a one-pixel band of “change” along its whole length, and a channel that shifted leavesTrees -> Wateron one bank andWater -> Treeson the other. Both concentrate on the long linear boundaries riparian work is about. The function adds eight columns to the patches fromdft_transition_vectors()— effective width2 * area / perimeter(width_m,width_px,flag_sliver), the share of a patch adjacent to the from-epoch interface with its to-class (boundary_frac,flag_boundary), and the nearest reverse-transition partner of comparable area (reciprocal_id,reciprocal_dist_m,flag_reciprocal) — and lets the caller decide. No composed verdict is returned, because real changes are sometimes genuinely thin; the docs show the one-liner. Stable (A -> A) rows get width only andNAelsewhere, meaning not applicable, not “checked clean”. -
Why
patch_area_mincannot do this. Area is width times length, and the artifact is pathological in width: a one-pixel band along a 5 km bank at 30 m is 15 ha, through any sane area threshold, while a real 0.5 ha cutblock is discarded as speckle. A 15 ha sliver and a 15 ha clearing are identical on the area axis and ~80x apart on the width axis. Width is orthogonal and composable — with it in hand,patch_area_mincan be lowered to recover small real changes without readmitting the slivers. The boundary signature is the discriminator width alone lacks: a registration artifact traces a boundary that already existed, a new road or harvested strip cuts across one and scores near zero. -
Reciprocity is proximity, not adjacency — a deliberate deviation from the issue. The two bands of a shifted channel sit on opposite banks with the stable channel between them, so they touch only when the river is one pixel wide (which is why the issue’s own probe found no touching pairs). The partner search is
sf::st_is_within_distance()atreciprocal_dist_maxpixels (default 5, a 50 m channel at 10 m) withmin(area) / max(area) >= reciprocal_area_ratio_min(default 0.5). -
Measured on the whole BULK floodplain (
bulk_co_ff04from stac-floodplains-bc: 386 km², a 169M-cell grid at 10 m, 2017 -> 2023): 21,701 change patches, 4,625 ha. 90% are slivers holding 23% of the change area; 68% sit on a pre-existing boundary (17.5% of area); 11% have a reciprocal partner within 5 px (5.4%), every pair symmetric, partner distances 0–50 m with a median of 22 m. The area filter does not remove them: atpatch_area_min = 5000, 165 boundary-hugging slivers totalling 123 ha survive, and only a 2 ha filter removes them all, at the cost of 48% of all mapped change. The artifact inflates gross change, not net: 15.7% of Trees loss and 26.4% of Trees gain carry the signature, yet net Trees moves only -867.5 -> -857.9 ha, because compensating shifts cancel by construction. Anyone reporting gross loss or gain from the transition table is carrying it. -
That run found a scale bug the unit tests structurally could not. The first boundary implementation held one in-memory 169M-cell raster per to-class per intermediate (1.35 GB each) — 11.3 GB peak for a single class, killed for memory at eight classes on a 64 GB machine. Every intermediate is now written to a temp file (
app,rasterize,segregate,focalwithfilename =) so terra streams in chunks, and onesegregate()stack with a singlefocal()and a singlezonal()replaces the per-class loop. The same run completes in 143 s with the pipeline’s peak memory set by the upstream transition stage. Unit pins were unchanged by the rewrite. -
sf::st_perimeter()is not used, on purpose. On projected data it delegates to lwgeom, which is not a dependency of sf or of drift — an undeclared dependencyR CMD checkdoes not report and a suite cannot see when the dev machine has it installed. Perimeter isst_length(st_boundary()), measured identical (max abs diff 0, holes and multipolygons included), with a test asserting lwgeom is never loaded. Barest_length()on polygons returns 0. Filed as soul#189. - New Geometric Artifacts section in the land-cover vignette, beside
patch_area_min, with the trajectory vignette named as the independent spectral confirmation route: that check confirms or refutes individual patches but needs a Sentinel-2 cube, cannot separate riparian deciduous trees from grass in peak summer, and cannot see the reciprocal relationship at all. The two are complements.
drift 0.12.0
-
Bug fix: the cache key moved on an rlang upgrade, silently orphaning every cached cube and fetch (#48). The key is the cache filename (
<year>_<key>.nc,cube_<key>.tif), andrlang::hash()was rewritten in rlang 1.3.0. Its own NEWS says so: “with this version all hash values will now be different… you should assume it’s always possible for a new version to invalidate existing hashes.” So every cache entry became unreachable — no error, no warning, no log line; the only symptom was “the pipeline got slower”, which is exactly the kind of thing nobody files. The frozen goldens were the only reason it was known. -
The reported diagnosis was wrong, and the correction changed the fix. The issue suspected sf/PROJ drift under
sf::st_as_binary(). Measured: the key function extracted from its own pinning commit reproduces today’s value exactly, so nothing in drift moved. Both pre-rlang-1.3.0 goldens fail to reproduce and the one re-pinned after the upgrade holds — and rlang 1.3.0 was installed between them. A fix aimed at the geometry member would have left the real cause untouched. -
Keys are now a function of content alone. Each member is rendered to a canonical string (geometry WKB as hex, numerics as IEEE-754 bytes, characters through
enc2utf8(), a one-character type tag per member) and that string is hashed by its bytes withdigest::digest(algo = "xxhash64", serialize = FALSE)— no R serializer, no library traversal of an R object.digestjoins Imports. It is a categorically different bet from rlang: it implements a published algorithm and pins this exact call shape to the upstream XXH64 reference vector in its own test suite, where rlang explicitly reserves the right to change. drift pins that same vector as a control test, so a future failure distinguishes “the hashing layer moved” from “our inputs changed” — a distinction the goldens alone could not make. - No rlang pin is needed: rlang is no longer in the key path at all.
-
The key is now 16 characters, not 12, which changes the documented
attr(, "cache_key")return value. A key collision does not crash — it silently serves the wrong raster, and nothing downstream can detect it. 48 bits to 64 costs four characters of filename and buys a 65,536x margin, and re-keying was already being paid for, so this was the only moment it was free. -
Cache entries move under a scheme directory,
<cache>/v2/<source>/. A deliberate key change is now a migration rather than a silent orphaning: superseded generations stay findable,dft_cache_info()reports them (n_files_superseded,size_mb_superseded), anddft_cache_clear(scheme = "superseded")reclaims that space. Nothing is deleted automatically. - What this costs you once: every existing entry is superseded — measured on one real cache, 204 files / 1.09 GB. Re-fetch is ~10 s per entry and flat in entry size (0.05 / 0.12 / 0.36 MB all took ~10 s; the cost is STAC query, signing and COG opens, not bytes), and it happens on demand rather than all at once. The issue’s “10–30 min per cube” is real for Sentinel-2 cubes but was not a cost anyone was paying — that cache held no cube entries at all. It is the forward-looking reason this matters: the next rlang bump would have destroyed those silently.
- Cross-machine: the key is now identical on any machine, R version, rlang version and architecture, so a cache can be shared or copied between machines and stay valid. It does not avoid a first-run rebuild on a new machine — the cache is machine-local — and the goldens double as the check, since they are now portable facts rather than facts about one environment.
- Three canonicalization hazards are guarded because they were measured, not assumed:
digest(serialize = FALSE)silently hashes only the first element of a character vector (a length > 1 would collapse every key to one value),sf::st_as_binary()returns a list sois.raw()isFALSEon it, andis.logical()must be branched beforeis.numeric()orTRUErenders as1. Numerics use IEEE-754 bytes rather thansprintf("%.17g"), which keepsNaNandNA_real_apart (is.na(NaN)isTRUE) and avoids a libc call whose formatting is a platform variable — this package has no cross-platform CI, so these goldens are verified on one machine.
drift 0.11.0
-
Bug fix: an interrupted fetch left a corrupt cache entry that every later run trusted as a hit (#41).
dft_stac_fetch()wrote each year’s raster straight to its canonical cache path, and the cache gate was presence-only — so a process killed mid-download left a partial file under the name a later run reports ascached, returns, and then breaks on ([mask] rasterization failed). Recovery was manual deletion of the specific<year>_<hash>.nc. Cache entries are now written to a unique temp file in the same directory and renamed into place only after a complete, validated write, so a killed fetch leaves at most a*.tmp*orphan that the gate can never serve. Latent on small AOIs, which is why it surfaced on long large-floodplain runs. -
The same fix lands on
dft_stac_cube(), which had the identical defect and was not named in the issue. It is the more expensive path to lose — a cube entry is a multi-hour Sentinel-2 stream rather than a small annual raster. -
A second instance of the same bug, which the issue did not name:
force = TRUEdestroyed a good cache. It truncated the canonical file before writing, so an interrupted forced re-fetch lost an entry that had been perfectly valid — the exact situation in which someone reaches forforce. There is a regression test. - Concurrent fetches of the same key no longer interleave. Two sessions writing one canonical path previously produced a mixed file; a unique temp per writer makes it clean last-writer-wins.
-
Cached entries are validated before they are served. Three damage shapes were measured rather than assumed, and no single cheaper check sees all of them: a tail-truncated or zero-byte file makes
terra::rast()error; the shape in the issue’s own traceback opens with a warning and a degenerate geotransform; and a content-damaged file whose size is preserved opens silently, with correct dim/res/CRS, masks fine, and returns plausible-looking values (nNA = 0) — its damage is visible only in the warning raised during the pixel read. SotryCatch(terra::rast())alone passes the second and third, and a geometry check alone passes the third. A failing entry is re-fetched with a warning naming the reason, never an abort telling the user to delete a file by hand. - Two limits of that check are deliberate and documented at the call sites, so they are not later “tightened” by someone reading them as oversights. The open check catches errors only, never the multidim-API warning — that is a capability message, and failing on it would refuse healthy
.ncentries across the whole cache. The empty-CRS check fires only in conjunction with an identity geotransform, since an absent CRS alone is not proof of damage and the cost of a false refusal is a silent, permanent re-download. -
The read probe samples rather than proves, and says so. It reads one row, so interior damage that leaves the container walkable can pass it; the guarantee against partial entries is the atomic write. On the cube path the probe is free and strictly stronger:
cube_check_nonempty()already scans every pixel viaterra::global(), so wrapping that existing call in a warning handler gives a whole-file check at no added cost. Validity is checked before it, so a truncated file that happens to read all-NAis reported as corrupt rather than as “your AOI does not overlap the collection”. -
No cache-format break. Unlike #51, nothing needs invalidating: all 168 entries in a real cache were checked and none is corrupt. The same 168 (123
.nc, 45.tif) are the false-refusal control for the new validator — it rejects none of them, at 41 ms each. -
dft_stac_fetch()no longer strands per-tile files whenmosaic_tiles()errors, and the GDAL PAM sidecar is carried across the rename (terra::writeRaster()emits one writing.nc, though not.tif), so it cannot be left behind under a dead temp name.
drift 0.10.0
-
Bug fix:
dft_stac_fetch()read only the first page of STAC items, so a wide AOI could silently build a raster from a truncated item set (#51). It calledrstac::get_request()with norstac::items_fetch(), and the partial item collection went straight togdalcubes::stac_image_collection()— no error, no warning, a plausible-looking raster with missing tiles over whatever the missing items covered. Because it depends on AOI size it would have appeared first on the largest, most-published area rather than in a test. The siblingdft_stac_cube()has always paged correctly;dft_stac_fetch()was simply never brought across. Fetch now pages to exhaustion through a new internalstac_items_paged(), signing after paging (signing first leaves every item from page 2 onward unsigned). -
The obvious guard for this is wrong, and drift deliberately does not implement it. Erroring when a
rel="next"link survives — what the issue proposed — would abort every correctly-paged fetch:rstac:::items_fetch.doc_itemsmutates onlyitems$featuresand neveritems$links, so a fully-paged collection still carries page 1’snext. Measured on the packaged AOI (io-lulc-annual-v02, 14 items ground truth): atlimit=1andlimit=3the fetch returns the complete 14 items and still reports anextlink, while atlimit=NULLandlimit=500it returns 14 with none. The link is stale, so it is now stripped beforeattr(result, "stac_items")reaches callers — otherwise anyone re-runningitems_fetch()on that attribute re-fetches pages 2..N into an already-complete feature list and silently duplicates them. Recorded ininst/notes/gdalcubes-pc-gotchas.mdso it is not re-litigated. - The strip matches
nextexactly, asrstacdoes, and a case-variant is deliberately kept rather than removed.rstac:::items_next.doc_itemsselects withlinks(items, rel == "next"), which is case-sensitive (measured:nextmatches,NEXTandNextdo not). So aNEXTlink is inert toitems_fetch()and cannot cause the duplicate it would otherwise be stripped to prevent — what it means instead is that rstac could not follow it and stopped after page one, which is the very truncation this release fixes. On Planetary Computer nothing else can detect that (items_matched()is NULL and duplicate ids cannot see a short read), so that link is the last local evidence.dft_stac_fetch()now warns and leaves it attached. An earlier draft of this fix stripped it case-insensitively, which would have deleted the evidence. - An item with no usable
idis now its own error rather than being folded into the duplicate check, which used to report two id-less items asduplicate item id: NA — pages overlapped— naming a cause that had not occurred. - Two completeness checks with deliberately different reach: duplicate item ids (never skipped, and the only signal available on Planetary Computer —
gdalcubes::stac_image_collection()drops duplicates behind a debug-only message, so nothing downstream would ever report them), anditems_matched()against the item count (fires only where the API reports a total; PC sends nonumberMatched, so this never executes there and is fixtured in the tests rather than left as dead code).limit = 500is a round-trip reducer, not the fix —items_fetch()is. -
Cache-format break: existing
dft_stac_fetch()caches rebuild once. The paging fix changes no cache-key parameter, so without a deliberate break a raster written from a truncated item set would keep being served by thefile.exists()short-circuit — and the wide-AOI users the bug hit hardest would get no fix at all on upgrade, silently and permanently underforce = FALSE.stac_cache_key()therefore gains a constant salt and the frozen-hash guardian moves from79f67b7b9daeto2264b5dbef6e. The cost is a one-time re-fetch of small annual land-cover rasters, not the multi-hour Sentinel-2 stream — which is whydft_stac_cube()chose a read-path check for its analogous problem and fetch can afford a key break.dft_stac_cube()caches were unaffected by #51 —stac_cube_cache_key()is a separate function and was untouched there. Superseded by #48 in 0.12.0, which re-keys both functions and moves every entry under av2/scheme directory; cube caches are invalidated too. -
dft_stac_fetch()now attachesattr(, "cache_key")alongside the existingattr(, "stac_items"), so a caller can record which cache entry served a fetch. It is per call, not per year — cached files are named<year>_<cache_key>, so one key covers every year the call returned. - Note for anyone reading warnings after upgrading:
gdalcubes::stac_image_collection()skips an unreadable item with a warning, so a now-complete (larger) item set can surface warnings that the truncated set never reached. That is the paging working, not a regression introduced here.
drift 0.9.0
-
dft_stac_cube()gainsparallel(defaultNULL=min(4, cores - 1)), and is roughly 2× faster out of the box on a machine with 5 or more cores as a result (#47; the auto value ismin(4, cores - 1), so a 1- or 2-core machine resolves to 1 and is unchanged). drift never calledgdalcubes::gdalcubes_options()anywhere, so every cube read has run single-threaded since the function existed — not a considered choice, just the consequence of never setting it. (gdalcubes also derives its default chunk size fromparallel, so raising it changes chunking as a side effect — but note that on this AOI finer chunking measured worse in isolation, 343.7 s / 693 requests at 128 px against 236.9 s / 462 at the default, so the speedup here is attributable to concurrency, not to chunk size.) Measured on the packaged AOI (4-month monthly kNDVI): 236.8 s atparallel = 1, 115.8 s at 4, 96.0 s at 8. Output is byte-identical at every setting — same grid, correlation 1.000, max absolute difference 0, same 49,244 non-NA cells — soparallelis a pure cost knob and deliberately does not enter the cache key; existing cached cubes stay valid. Capped at 4 by default rather than the full core count because each gdalcubes worker holds chunks in memory and drift has hit OOM on large AOIs before (#27, #34); raise it where there is headroom, or pass1for the previous behaviour. -
gdalcubes::filter_geom()was evaluated and rejected on measurement (#47, #38). The segfault that originally blocked it is fixed inNewGraphEnvironment/gdalcubes@newgraph, so it was available to adopt; it is simply not worth it. It skips whole chunks rather than pixels, and gdalcubes clamps a chunk edge to[64, 1024]px — coarse enough to skip nothing at the default (462 range requests and 236.9 s, against the bbox baseline’s 462 / 236.8 s), and fine enough to break alignment with the COGs’ 512×512 blocks when forced down (693 requests / 348.2 s at 64 px — 1.5× the requests and 47% slower, to skip 26.7% of the ground). The AOI/bbox area ratio the proposal rested on is the wrong bound: the read is chunk-granular and the cost is COG-block-granular. drift therefore keeps itsterra::mask()clip, and the reasoning is recorded ininst/notes/gdalcubes-pc-gotchas.mdso it is not re-litigated. -
Bug fix: the end-to-end cube test passed on an all-
NAcube. Its coverage assertion wasmean(rowSums(!is.na(values)) == 0) > 0.5, which an entirely empty cube satisfies with1.0— as it did the layer-count, time and cache-file assertions. Every assertion in drift’s only network cube test was satisfied by a cube containing no data, i.e. by exactly the gdalcubes failure mode #32 exists to prevent. Replaced with in-polygon non-NA, per-layer non-NA, and outside-NAassertions, with the in-polygon cell set derived from the AOI geometry rather than from the cube’s own values.dft_stac_cube()also now aborts, before writing anything to the cache, when the assembled cube has no data on any layer (deliberately the whole stack, not the first layer: an individual layer being empty is documentedmonthsbehaviour, so a first-layer check would abort the package’s ownmonths = 6:9example after the full COG stream) — so an empty cube can no longer be cached and served for every later call. -
Documentation fix:
stac_cube_clip()documented the wrong clip rule. It claimed cells whose centre falls outside the polygon becomeNA;terra::mask()defaults totouches = TRUE, so the clip is inclusive at the boundary by up to one cell. The difference is 15.5% of the analysed footprint on the packaged AOI, which matters to anyone reasoning about boundary hectares. The rule is now pinned by a test using an irregular polygon — the existing test could not catch it, because its polygon is axis-aligned on a cell boundary where both rules agree. -
dft_stac_cube()’stile_sizedocumentation now carries its measured cost: at 640 m on the packaged AOI it is the slowest option (1263.6 s / 3213 requests against an untiled 236.8 s / 462), because every tile rebuilds the image collection and reopens the COGs. Preferparallelfor speed; reach fortile_sizeonly when peak memory rather than wall clock is the constraint.
drift 0.8.0
-
dft_transition_attribute()tags change patches fromdft_transition_vectors()with columns from any overlay polygon layer — fire perimeters, cutblocks, roads, tenures — so a driver can separate mapped transitions by cause without hand-rolling spatial joins (#42). Deliberately generic: drift carries no BC/domain knowledge; the caller supplies the overlay, the columns to carry (cols), and optionally a numeric temporal filter (time_col+time_interval, bounds inclusive) that keeps only overlay features whose time falls within the transition interval — a 2022 fire attributes a 2017→2023 loss, a 2012 fire does not. Two assignment modes for a patch that straddles multiple overlay features:match_mode = "all"(left join, one row per match) or"largest"(one row per patch by greatest overlap area). Largest-overlap assignment is intersection-based, so combining it with a custompredicateis an error rather than a silent mis-attribution; the overlay is reprojected to the patch CRS and run throughst_make_valid()automatically, since real-world disturbance perimeters routinely fail validity checks.
drift 0.7.0
-
dft_stac_cube()gainstile_size(defaultNULL), the continuous-path twin ofdft_stac_fetch()’stile_size(#36): an opt-in that bounds the STAC read to the AOI footprint (#38). By default one gdalcubes cube is streamed over the whole AOI bounding box, so the COG streaming — the dominant cost (~10-30 min for a multi-year monthly Sentinel-2 fetch) — scales with the bbox, not the AOI; for a thin, diagonal floodplain corridor the bbox is largely empty. Whentile_size(CRS units — metres for the default UTM CRS) is set, the bbox is split into ares-aligned grid and only tiles that intersect the AOI polygon are streamed — each carrying the full SCL mask, spectral index, and 2022 baseline-offset split — then mosaicked withterra::merge(), so a corridor reads close to its footprint. This is thefilter_geom-independent path (the polygon clip that would do this in-cube segfaults on the pinned gdalcubes build). The cube always caches a.tifand a tiled read keys distinctly, so untiled caches are untouched andtile_size = NULLis byte-for-byte the previous behavior. Because the cube resamples with bilinear, a tiled cube faithfully reproduces the untiled cube (bilinear-aligned correlation ~0.997, per-layer means within ~1e-3, no tile seams) but lands on a bbox-anchored grid that is sub-pixel-offset from — not pixel-identical to — the untiled cube; the offset is immaterial to the per-pixeldft_rast_break()/dft_rast_trend()reducers. Withtile_sizeset,clip = FALSEreturns the AOI-intersecting tile union (withNAwhere empty tiles were skipped), not a gap-free bounding box.
drift 0.6.0
-
dft_stac_fetch()gainstile_size(defaultNULL), an opt-in that bounds the STAC download to the AOI footprint (#36). By default a single cube is streamed over the whole AOI bounding box, so for a thin, diagonal floodplain corridor (measured ~10% of the bbox inside the polygon) roughly 10× more pixels are downloaded than the AOI needs. Whentile_size(CRS units — metres for the default UTM CRS) is set, the bbox is split into ares-aligned grid and only tiles that intersect the AOI polygon are streamed, then mosaicked withterra::merge()— so a corridor fetches close to its footprint. Smaller tiles waste less bbox but cost more per-tile round trips (no auto-tuning). This is thefilter_geom-independent path (the polygon-clip that would do this in the cube pipeline segfaults on the pinned gdalcubes build). Tiled fetches cache a terra GeoTIFF (.tif) rather than a gdalcubes NetCDF (.nc) and key distinctly, so existing untiled caches are untouched;tile_size = NULLis byte-for-byte the previous behavior. The same read residual on the continuousdft_stac_cube()path is tracked as #38.
drift 0.5.0
-
dft_stac_cube()gainsclip(defaultTRUE), restoring AOI-polygon-tight output (#32). The assembled index stack is masked to the AOI polygon withterra::mask()— client-side, becausegdalcubes::filter_geom()segfaults / returns an all-NA cube on the pinned build — so cells outside the polygon areNAon every layer. The reduced raster fromdft_rast_break()/dft_rast_trend()is now polygon-tight with no caller-side mask, and those reducers skip out-of-AOI pixels via their valid-observation gate.clip = FALSEkeeps the full bounding box. This is an output change for callers that relied on the bounding-box extent, and the clip is folded into the cube cache key, so existing cached cubes rebuild once. Note the clip affects the output only — the full bbox of COGs is still streamed either way (the AOI cannot be pushed into the read on the pinned gdalcubes build).
drift 0.4.0
- Categorical land-cover change detection no longer exhausts memory on large-floodplain AOIs (#34, #28).
dft_rast_transition()was rewritten to stream entirely throughterra— transitions are encoded and filtered with raster arithmetic,terra::subst(),patches(), and a singleterra::freq(), with noterra::values()pull and no full-grid R vectors — so peak memory scales with the number of distinct transitions and patches, not the grid size (producer-only peak at 16M cells dropped from 2.66 GB to 1.63 GB). Output is byte-identical to the previous version, verified by a golden snapshot across the full parameter matrix. -
dft_transition_vectors()gainschanges_only(defaultFALSE): whenTRUE, stable (from == to) transitions are dropped at the raster level before polygonizing, soterra::as.polygons()only builds geometry for actual change patches. On a fragmented floodplain — where the stable mosaic is most of the grid and polygonization dominates memory — this roughly halves peak use (a 9M-cell, 415k-patch benchmark went from 3.83 GB to 1.71 GB). The result equals the default output filtered to change patches. Whenpatch_area_minis set, small patches are also dropped before polygonizing, with identical output. -
patch_idindft_transition_vectors()is numbered over the surviving patches when filtering drops any, and an empty result now carries the zone column so per-zone results bind cleanly.
drift 0.3.0
- Continuous index-trajectory change detection for floodplain reaches (#30). A new fetch-and-reduce pipeline complements the categorical
dft_stac_fetch()path.dft_stac_cube()builds a cloud-masked monthly spectral-index stack from Sentinel-2 (via a new"sentinel-2-l2a"source);dft_rast_break()reduces it per pixel withbfast::bfastmonitor()into a two-band raster of abrupt break date and magnitude; anddft_rast_trend()reduces it to a per-pixel gradual trend — a robust Theil-Sen slope (index change per year) with Mann-Kendall significance — for degradation/recovery monitoring the annual labels cannot show. Together they let a continuous trajectory validate categorical land-cover transitions (confirming which mapped losses carry a real spectral decline) and detect gradual change. See the “Trajectories as a Check on Land-Cover Change” vignette. -
dft_index_expr()anddft_index_table()add a table-driven spectral-index registry (NDVI, kNDVI, NDMI) whose formulas are written over band roles, so one index resolves against any reflectance source; the reflectance scale/offset is folded into each expression. - Sentinel-2 handling is correctness-focused:
dft_stac_cube()masks cloud/shadow/cirrus/snow, restricts to caller-chosen calendarmonths(e.g. the growing season) to sharpen the signal and cut scenes streamed, and — because the +1000 DN reflectance offset only applies from processing baseline 04.00 (2022-01-25) — splits items at that boundary and corrects each side, so a multi-year series carries no artificial index step at 2022. -
dft_stac_config()gains a role-based schema for reflectance cube sources (band roles, mask classes, scale/offset, offset boundary), leaving the categoricalio-lulc/esa-worldcoversources unchanged.bfastadded to Suggests. - Known limitation tracked as a follow-up: the cube spans the AOI bounding box rather than the polygon (a gdalcubes
filter_geomlimitation, #32); labelling breaks with from/to land-cover classes is #31.
drift 0.2.4
-
dft_transition_vectors()no longer exhausts memory on large-extent rasters (#27). The per-class loop allocated full-grid vectors per class and per patch — ncell × n_patches churn that OOM-killed a 102.6M-cell, 56-class floodplain. Replaced by a singleterra::patches(values = TRUE)pass plus a sparse patch-to-label map. Output is identical (verified patch-by-patch against the old implementation); onlypatch_idnumbering / row order changes, to raster scan order. Benchmark at 24M cells: 1.9 s for a 4,799-patch raster; the old code took 122 s on a milder 1,232-patch raster of the same size. - terra dependency floored at
>= 1.8-10: earlier versions had an edge-wraparound bug inpatches(values = TRUE)that silently merged patches touching opposite raster edges.
drift 0.2.3
- Fix silent cross-AOI cache collision in
dft_stac_fetch()(#25). Cache files were keyed by source + year only, so fetching a second AOI with the same source/year silently returned the first AOI’s raster masked to the second AOI’s extent. Cache filenames now include a hash of the AOI geometry and all fetch-affecting parameters (res,crs,dt,aggregation,resampling,stac_url,collection,asset). Existing caches re-fetch on first use after upgrading;dft_cache_clear()reclaims the orphaned old-format files. -
force = TRUEnow overwrites the cached file instead of erroring with “File already exists” (#25).
drift 0.2.2
- Startup quote pool expanded to 113. Adds 52 domain-expert quotes from 11 voices across floodplain/river process (David Montgomery, Ellen Wohl), Indigenous stewardship (Robin Wall Kimmerer, Kyle Whyte, Nancy Turner, Jeannette Armstrong), ecosystem valuation (Kai Chan), Canadian public voices (David Suzuki, Wade Davis), and legacy conservation (Aldo Leopold, Wendell Berry).
- Tim Beechie was on the target list but yielded zero — no public interview / podcast / documentary footprint. Process-paper voice only.
- Same rigor as v0.2.1: parallel research agents, independent fact-check pass (3 dropped for misattribution or text drift, 2 fixed from fact-check flags).
drift 0.2.1
- Startup quote ritual:
library(drift)prints a random fact-checked quote from 15 hip-hop artists on attach. Italic quote, grey attribution, clickable bluesourcehyperlink to the primary-source interview. Suppress viaoptions(drift.quote_show_source = FALSE). - Curated via the soul
/quotes-enableskill using multi-agent research + independent primary-source fact-check. 61 entries. Seedata-raw/quotes_build.Rfor full provenance. -
cliadded to Imports for OSC 8 hyperlinks and styling inR/zzz.R.
drift 0.2.0
-
dft_rast_transition()— addpatch_area_minparameter to filter small connected patches of changed pixels; return$removedraster for visual QA of filtered patches; addfrom_class/to_classfilters -
dft_transition_vectors()— vectorize transition raster into sf polygons with per-patch area, transition labels, and optional zone attribution -
dft_rast_consensus()— per-pixel mode across classified rasters for temporal noise filtering; optional confidence layer -
dft_map_interactive()— newtransitionparameter overlays transition layers as checkboxes; Google Satellite and Esri Satellite basemaps; custom tile URL support -
dft_check_crs()— internal helper that errors on geographic CRS input; wired intodft_rast_transition()anddft_rast_summarize() - Vignette: transition detection, tree loss filtering, patch area filtering with comparison table, interactive map with transition overlays
drift 0.1.0
Initial public release.
-
dft_stac_fetch()— fetch classified rasters from STAC catalogs via gdalcubes -
dft_rast_classify()— apply class labels, colors, and optional remap to SpatRasters -
dft_rast_summarize()— compute area by class and year with unit conversion -
dft_map_interactive()— interactive leaflet map with layer toggle, legend, fullscreen, and titiler COG support -
dft_class_table()— shipped class tables for IO LULC and ESA WorldCover -
dft_stac_config()— STAC endpoint registry - Cache management:
dft_cache_path(),dft_cache_info(),dft_cache_clear() - Vignette: Neexdzii Kwa floodplain land cover change 2017-2023
