Skip to content

pysmo.tools.cache

On-disk waveform caching: fetch a station and window once, replay it later.

FetchCache is the main entry point. Built from a fetch function and a parser, it behaves as a callable (station, starttime, endtime) -> Seismogram. The first call for a given station and window downloads and stores the raw response; later calls for the same station and window read it back from the file, with no network access. Its call signature is SeismogramFetcher, so it can also be a PysmoProject's fetch_seismogram.

BlobCache is the storage layer underneath: a keyed store of byte blobs in a single SQLite file, compressed, with an optional size cap. FetchCache uses it for raw fetch responses; TransformCache uses it for transformed seismograms. Use it directly to cache anything else that is costly to produce and reduces to bytes.

Portable on disk

A cache file is an ordinary SQLite database: one table, each row a zlib-compressed entry. A SQLite client and zlib are all it takes to read the contents back, in any language.

Examples:

Fetch a window of SAC data once, then replay it from disk on the next run. FetchCache is paired here with fetch_sac and SAC.from_zip.

>>> import pandas as pd
>>> from pysmo import MiniStation, Seismogram
>>> from pysmo.classes import SAC
>>> from pysmo.tools.cache import FetchCache
>>> from pysmo.tools.web import fetch_sac
>>>
>>> def parse_sac_seismogram_zip(raw: bytes) -> Seismogram:
...     return SAC.from_zip(raw).seismogram
...
>>> station = MiniStation(
...     name="ANMO", network="IU", location="00", channel="LHZ",
...     latitude=34.945981, longitude=-106.457133,
... )
>>> starttime = pd.Timestamp("2010-02-27T06:44:00Z")
>>> endtime = pd.Timestamp("2010-02-27T06:54:00Z")
>>>
>>> cache = FetchCache(
...     path="waveform_cache.sqlite3", fetch_raw=fetch_sac, parse=parse_sac_seismogram_zip
... )
>>> seismogram = cache(station, starttime, endtime)  # miss: fetches and stores
>>> seismogram_again = cache(station, starttime, endtime)  # hit: no fetch
>>> seismogram_again.data.shape == seismogram.data.shape
True
>>>

Type Aliases:

Name Description
RawParser

A callable that parses raw fetch bytes into a Seismogram.

Classes:

Name Description
BlobCache

A keyed store of byte blobs in a single SQLite file.

FetchCache

A waveform cache: fetch a station and window once, re-parse it from a file.

RawFetcher

A callable that returns the raw bytes for a station and time window.

RawParser

RawParser = Callable[[bytes], Seismogram]

A callable that parses raw fetch bytes into a Seismogram.

For example SAC.from_zip or MSeed.from_bytes. Must match the format the RawFetcher it is paired with returns.

BlobCache

A keyed store of byte blobs in a single SQLite file.

get takes a string key and a callback. On a hit it returns the stored blob; on a miss it calls the callback, stores what it returns (compressed), and returns that. peek and put are the same two halves on their own. Keys and values are arbitrary bytes; the cache interprets neither.

Pass max_bytes to cap the total stored size; once it is exceeded entries are removed in insertion order (not least-recently-used) until the cache fits again.

Local disk only

The SQLite file must be on local disk. Several processes on one machine may share it — writers serialise on SQLite's own file lock — but WAL mode and access over a network filesystem are unsupported and can corrupt the file. Within a process, concurrent get calls for the same missing key each run the callback and return their own result; the first write is the one kept.

Examples:

>>> from pathlib import Path
>>> import tempfile
>>> from pysmo.tools.cache import BlobCache
>>>
>>> tmp = Path(tempfile.mkdtemp())
>>> cache = BlobCache(path=tmp / "blobs.sqlite3", encoding_version=1)
>>> calls = []
>>> def produce() -> bytes:
...     calls.append(1)
...     return b"payload"
...
>>> cache.get("some-key", produce)
b'payload'
>>> cache.get("some-key", produce)  # hit: produce not called again
b'payload'
>>> len(calls)
1
>>>

Methods:

Name Description
__attrs_post_init__

Fail fast if path's parent directory doesn't exist.

__del__

Close the connection when the cache is garbage collected.

__getstate__

Drop the live connection and lock for pickling.

__setstate__

Restore the fields and create a fresh lock.

close

Close the database connection.

get

Return the blob stored under key, producing and storing it on a miss.

peek

Return the blob stored under key, or None if there is none.

put

Store value under key (compressed), evicting if over max_bytes.

Attributes:

Name Type Description
encoding_version PositiveInt

Layout version for the file, recorded on creation and checked on open.

max_bytes PositiveInt | None

Cap on the total compressed size of stored blobs, in bytes.

path Path

Path to the SQLite file.

wal bool

Enable SQLite WAL mode (local disk only).

Source code in src/pysmo/tools/cache.py
@define(kw_only=True)
class BlobCache:
    """A keyed store of byte blobs in a single SQLite file.

    [`get`][pysmo.tools.cache.BlobCache.get] takes a string key and a
    callback. On a hit it returns the stored blob; on a miss it calls the
    callback, stores what it returns (compressed), and returns that.
    [`peek`][pysmo.tools.cache.BlobCache.peek] and
    [`put`][pysmo.tools.cache.BlobCache.put] are the same two halves on their
    own. Keys and values are arbitrary bytes; the cache interprets neither.

    Pass `max_bytes` to cap the total stored size; once it is exceeded entries
    are removed in insertion order (not least-recently-used) until the cache
    fits again.

    Warning: Local disk only
        The SQLite file must be on local disk. Several processes on one
        machine may share it — writers serialise on SQLite's own file lock —
        but WAL mode and access over a network filesystem are unsupported and
        can corrupt the file. Within a process, concurrent
        [`get`][pysmo.tools.cache.BlobCache.get] calls for the same missing
        key each run the callback and return their own result; the first
        write is the one kept.

    Examples:
        ```python
        >>> from pathlib import Path
        >>> import tempfile
        >>> from pysmo.tools.cache import BlobCache
        >>>
        >>> tmp = Path(tempfile.mkdtemp())
        >>> cache = BlobCache(path=tmp / "blobs.sqlite3", encoding_version=1)
        >>> calls = []
        >>> def produce() -> bytes:
        ...     calls.append(1)
        ...     return b"payload"
        ...
        >>> cache.get("some-key", produce)
        b'payload'
        >>> cache.get("some-key", produce)  # hit: produce not called again
        b'payload'
        >>> len(calls)
        1
        >>>
        ```
    """

    path: Path = field(converter=Path)
    """Path to the SQLite file.

    Created on first use; its parent directory must already exist.
    """

    encoding_version: PositiveInt
    """Layout version for the file, recorded on creation and checked on open.

    Opening a file that was written with a different value raises
    `ValueError`. Each cache built on `BlobCache` passes its own constant.
    """

    wal: bool = False
    """Enable SQLite WAL mode (local disk only)."""

    max_bytes: PositiveInt | None = field(
        default=None,
        validator=validators.optional(validators.gt(0)),
    )
    """Cap on the total compressed size of stored blobs, in bytes.

    When a new entry pushes the total past the cap, the oldest entries are
    removed until it fits again. A single entry larger than the cap is stored
    and kept anyway, so it is never re-produced on every call. `None` (the
    default) means no cap. The file on disk is somewhat larger than
    `max_bytes` because of SQLite's own page and index overhead.
    """

    _conn: sqlite3.Connection | None = field(
        init=False, default=None, repr=False, eq=False
    )
    _lock: threading.RLock = field(
        init=False, factory=threading.RLock, repr=False, eq=False
    )
    """Serialises connection setup, reads and writes; the cache may be called
    from more than one thread. Reentrant so a locked read or write can call
    `_connect` without releasing first."""

    def __attrs_post_init__(self) -> None:
        """Fail fast if `path`'s parent directory doesn't exist."""
        if not self.path.parent.is_dir():
            raise FileNotFoundError(
                f"Parent directory does not exist: {self.path.parent}"
            )

    def __getstate__(self) -> dict[str, Any]:
        """Drop the live connection and lock for pickling."""
        state = attrs_getstate(self, {"_conn": None})
        del state["_lock"]
        return state

    def __setstate__(self, state: dict[str, Any]) -> None:
        """Restore the fields and create a fresh lock."""
        attrs_setstate(self, state)
        object.__setattr__(self, "_lock", threading.RLock())

    def close(self) -> None:
        """Close the database connection.

        Optional: the connection is also closed when the cache is garbage
        collected. Call this to release the handle sooner.
        """
        with self._lock:
            if self._conn is not None:
                self._conn.close()
                self._conn = None

    def __del__(self) -> None:
        """Close the connection when the cache is garbage collected."""
        conn = getattr(self, "_conn", None)
        if conn is not None:
            conn.close()

    def _connect(self) -> sqlite3.Connection:
        with self._lock:
            if self._conn is None:
                conn = sqlite3.connect(self.path, timeout=30, check_same_thread=False)
                if self.wal:
                    conn.execute("PRAGMA journal_mode=WAL")
                version = conn.execute("PRAGMA user_version").fetchone()[0]
                if version == 0:
                    conn.execute(f"PRAGMA user_version = {self.encoding_version}")
                elif version != self.encoding_version:
                    conn.close()
                    raise ValueError(
                        f"{self.path} was written with a different cache "
                        + f"encoding (user_version={version}, expected "
                        + f"{self.encoding_version})."
                    )
                conn.execute(_CREATE_CACHE_TABLE)
                conn.execute(_CREATE_CACHE_STATS_TABLE)
                conn.execute(_CREATE_CACHE_INSERT_TRIGGER)
                conn.execute(_CREATE_CACHE_DELETE_TRIGGER)
                row = conn.execute(
                    "SELECT total_bytes FROM cache_stats WHERE id = 1"
                ).fetchone()
                if row is None:
                    conn.execute(
                        "INSERT OR IGNORE INTO cache_stats (id, total_bytes) "
                        + "VALUES (1, (SELECT COALESCE(SUM(length(data)), 0) FROM cache))"
                    )
                # Commit the schema and seeded counter now, so a session that
                # only ever reads (never reaching `_store`'s commit) doesn't
                # roll this setup back on close and redo it next run.
                conn.commit()
                self._conn = conn
            return self._conn

    def get(self, key: str, produce: Callable[[], bytes]) -> bytes:
        """Return the blob stored under `key`, producing and storing it on a miss.

        `produce` runs outside the lock, so a slow callback does not block
        other threads; two of them racing the same missing key both run it.

        Args:
            key: The cache key.
            produce: Called only on a miss. Must return `bytes`, which are
                stored compressed.

        Returns:
            The blob: read from the file on a hit, from `produce` on a miss.
        """
        hit = self.peek(key)
        if hit is not None:
            return hit
        produced = produce()
        self.put(key, produced)
        return produced

    def peek(self, key: str) -> bytes | None:
        """Return the blob stored under `key`, or `None` if there is none.

        Never runs a callback. Raises `ValueError` if the stored blob is
        corrupt or decompresses past an internal size ceiling.
        """
        with self._lock:
            conn = self._connect()
            row = conn.execute(
                "SELECT data FROM cache WHERE key = ?", (key,)
            ).fetchone()
        return _decompress(row[0]) if row is not None else None

    def put(self, key: str, value: bytes) -> None:
        """Store `value` under `key` (compressed), evicting if over `max_bytes`.

        A first write for a key wins; a later `put` for the same key is
        ignored. A failure to write to the file is warned about, not raised —
        the caller keeps whatever it was about to cache.
        """
        compressed = zlib.compress(value)
        with self._lock:
            conn = self._connect()
            try:
                self._store(conn, key, compressed)
            except sqlite3.Error as exc:
                warnings.warn(
                    f"could not write to cache {self.path}: {exc}",
                    stacklevel=2,
                )

    def _store(self, conn: sqlite3.Connection, key: str, compressed: bytes) -> None:
        with self._lock, conn:
            cursor = conn.execute(
                "INSERT OR IGNORE INTO cache (key, data) VALUES (?, ?)",
                (key, compressed),
            )
            if cursor.rowcount > 0:
                if cursor.lastrowid is None:
                    raise RuntimeError("newly inserted cache row has no rowid")
                self._evict(conn, cursor.lastrowid)

    def _evict(self, conn: sqlite3.Connection, inserted_rowid: int) -> None:
        max_bytes = self.max_bytes
        if max_bytes is None:
            return
        if self._tracked_total(conn) <= max_bytes:
            return
        real = self._real_total(conn)
        if real <= max_bytes:
            # Tracked total overstated real usage (e.g. an uncounted external
            # delete); heal the counter without evicting anything.
            conn.execute("UPDATE cache_stats SET total_bytes = ? WHERE id = 1", (real,))
            return
        # Truly over limit: evict the oldest entries down to the low-water
        # mark, so a full cache does not re-check ground truth every miss.
        target = int(max_bytes * _LOW_WATER_FRACTION)
        freed = self._delete_oldest(conn, inserted_rowid, real - target)
        conn.execute(
            "UPDATE cache_stats SET total_bytes = ? WHERE id = 1", (real - freed,)
        )

    def _tracked_total(self, conn: sqlite3.Connection) -> int:
        row = conn.execute(
            "SELECT total_bytes FROM cache_stats WHERE id = 1"
        ).fetchone()
        if row is not None and row[0] >= 0:
            return row[0]
        # Rebuild if the cache_stats row is missing or stale. A negative total
        # is never legitimate and means the counter no longer matches which
        # rows the triggers have accounted for.
        real = self._real_total(conn)
        conn.execute(
            "INSERT OR REPLACE INTO cache_stats (id, total_bytes) VALUES (1, ?)",
            (real,),
        )
        return real

    def _real_total(self, conn: sqlite3.Connection) -> int:
        return int(
            conn.execute("SELECT COALESCE(SUM(length(data)), 0) FROM cache").fetchone()[
                0
            ]
        )

    def _delete_oldest(
        self, conn: sqlite3.Connection, inserted_rowid: int, excess: int
    ) -> int:
        cursor_iter = conn.execute(
            "SELECT rowid, length(data) FROM cache "
            + "WHERE rowid != ? "
            + "ORDER BY rowid ASC",
            (inserted_rowid,),
        )
        to_delete: list[int] = []
        freed = 0
        for rowid, length in cursor_iter:
            to_delete.append(rowid)
            freed += length
            if freed >= excess:
                break
        cursor_iter.close()
        for chunk in batched(to_delete, 500):
            placeholders = ",".join("?" * len(chunk))
            conn.execute(f"DELETE FROM cache WHERE rowid IN ({placeholders})", chunk)
        return freed

encoding_version instance-attribute

encoding_version: PositiveInt

Layout version for the file, recorded on creation and checked on open.

Opening a file that was written with a different value raises ValueError. Each cache built on BlobCache passes its own constant.

max_bytes class-attribute instance-attribute

max_bytes: PositiveInt | None = field(
    default=None,
    validator=validators.optional(validators.gt(0)),
)

Cap on the total compressed size of stored blobs, in bytes.

When a new entry pushes the total past the cap, the oldest entries are removed until it fits again. A single entry larger than the cap is stored and kept anyway, so it is never re-produced on every call. None (the default) means no cap. The file on disk is somewhat larger than max_bytes because of SQLite's own page and index overhead.

path class-attribute instance-attribute

path: Path = field(converter=Path)

Path to the SQLite file.

Created on first use; its parent directory must already exist.

wal class-attribute instance-attribute

wal: bool = False

Enable SQLite WAL mode (local disk only).

__attrs_post_init__

__attrs_post_init__() -> None

Fail fast if path's parent directory doesn't exist.

Source code in src/pysmo/tools/cache.py
def __attrs_post_init__(self) -> None:
    """Fail fast if `path`'s parent directory doesn't exist."""
    if not self.path.parent.is_dir():
        raise FileNotFoundError(
            f"Parent directory does not exist: {self.path.parent}"
        )

__del__

__del__() -> None

Close the connection when the cache is garbage collected.

Source code in src/pysmo/tools/cache.py
def __del__(self) -> None:
    """Close the connection when the cache is garbage collected."""
    conn = getattr(self, "_conn", None)
    if conn is not None:
        conn.close()

__getstate__

__getstate__() -> dict[str, Any]

Drop the live connection and lock for pickling.

Source code in src/pysmo/tools/cache.py
def __getstate__(self) -> dict[str, Any]:
    """Drop the live connection and lock for pickling."""
    state = attrs_getstate(self, {"_conn": None})
    del state["_lock"]
    return state

__setstate__

__setstate__(state: dict[str, Any]) -> None

Restore the fields and create a fresh lock.

Source code in src/pysmo/tools/cache.py
def __setstate__(self, state: dict[str, Any]) -> None:
    """Restore the fields and create a fresh lock."""
    attrs_setstate(self, state)
    object.__setattr__(self, "_lock", threading.RLock())

close

close() -> None

Close the database connection.

Optional: the connection is also closed when the cache is garbage collected. Call this to release the handle sooner.

Source code in src/pysmo/tools/cache.py
def close(self) -> None:
    """Close the database connection.

    Optional: the connection is also closed when the cache is garbage
    collected. Call this to release the handle sooner.
    """
    with self._lock:
        if self._conn is not None:
            self._conn.close()
            self._conn = None

get

get(key: str, produce: Callable[[], bytes]) -> bytes

Return the blob stored under key, producing and storing it on a miss.

produce runs outside the lock, so a slow callback does not block other threads; two of them racing the same missing key both run it.

Parameters:

Name Type Description Default
key str

The cache key.

required
produce Callable[[], bytes]

Called only on a miss. Must return bytes, which are stored compressed.

required

Returns:

Type Description
bytes

The blob: read from the file on a hit, from produce on a miss.

Source code in src/pysmo/tools/cache.py
def get(self, key: str, produce: Callable[[], bytes]) -> bytes:
    """Return the blob stored under `key`, producing and storing it on a miss.

    `produce` runs outside the lock, so a slow callback does not block
    other threads; two of them racing the same missing key both run it.

    Args:
        key: The cache key.
        produce: Called only on a miss. Must return `bytes`, which are
            stored compressed.

    Returns:
        The blob: read from the file on a hit, from `produce` on a miss.
    """
    hit = self.peek(key)
    if hit is not None:
        return hit
    produced = produce()
    self.put(key, produced)
    return produced

peek

peek(key: str) -> bytes | None

Return the blob stored under key, or None if there is none.

Never runs a callback. Raises ValueError if the stored blob is corrupt or decompresses past an internal size ceiling.

Source code in src/pysmo/tools/cache.py
def peek(self, key: str) -> bytes | None:
    """Return the blob stored under `key`, or `None` if there is none.

    Never runs a callback. Raises `ValueError` if the stored blob is
    corrupt or decompresses past an internal size ceiling.
    """
    with self._lock:
        conn = self._connect()
        row = conn.execute(
            "SELECT data FROM cache WHERE key = ?", (key,)
        ).fetchone()
    return _decompress(row[0]) if row is not None else None

put

put(key: str, value: bytes) -> None

Store value under key (compressed), evicting if over max_bytes.

A first write for a key wins; a later put for the same key is ignored. A failure to write to the file is warned about, not raised — the caller keeps whatever it was about to cache.

Source code in src/pysmo/tools/cache.py
def put(self, key: str, value: bytes) -> None:
    """Store `value` under `key` (compressed), evicting if over `max_bytes`.

    A first write for a key wins; a later `put` for the same key is
    ignored. A failure to write to the file is warned about, not raised —
    the caller keeps whatever it was about to cache.
    """
    compressed = zlib.compress(value)
    with self._lock:
        conn = self._connect()
        try:
            self._store(conn, key, compressed)
        except sqlite3.Error as exc:
            warnings.warn(
                f"could not write to cache {self.path}: {exc}",
                stacklevel=2,
            )

FetchCache

A waveform cache: fetch a station and window once, re-parse it from a file.

Call it as (station, starttime, endtime) -> Seismogram. On the first call for a given station and window it runs fetch_raw, stores the raw response, and returns parse of it; later calls for the same station and window read the stored bytes and return parse of those, without fetching. parse therefore runs on every call, and the file holds the unparsed response, which any other tool can read. For a project that should also skip the transform on a hit, see TransformCache.

Any format works, as long as fetch_raw and parse agree (e.g. fetch_sac with SAC.from_zip). While a window stays cached it is replayed byte-for-byte; an entry evicted under a finite max_bytes is fetched again on next access.

A response is stored only once parse has accepted it and it is non-empty, so an empty body (an FDSN 204 for a gap) or a response that does not parse is retried on the next call rather than cached as a permanent failure. Two threads racing the same uncached window both fetch.

The call signature is SeismogramFetcher, so an instance also serves as a PysmoProject's fetch_seismogram. A database written by pysmo's earlier SqliteArchiveFetcher is read as-is: the constructor is the same, and existing .sqlite3 files need no migration.

Methods:

Name Description
__attrs_post_init__

Build the inner store (which also checks path's parent exists).

__call__

Return the Seismogram for station and window, from cache when possible.

__getstate__

Drop the inner store; it is rebuilt from the plain fields on unpickling.

__setstate__

Restore the plain fields, then rebuild the inner store.

close

Close the inner store's connection, if one is open.

Attributes:

Name Type Description
fetch_raw RawFetcher

Fetches the raw response for a station and time window.

max_bytes PositiveInt | None

Cap on the total stored size in bytes; see

parse RawParser

Parses a raw response into a Seismogram. Runs on every call.

path Path

Path to the SQLite file.

wal bool

Enable SQLite WAL mode (local disk only).

Source code in src/pysmo/tools/cache.py
@define(kw_only=True)
class FetchCache:
    """A waveform cache: fetch a station and window once, re-parse it from a file.

    Call it as `(station, starttime, endtime) -> Seismogram`. On the first
    call for a given station and window it runs `fetch_raw`, stores the raw
    response, and returns `parse` of it; later calls for the same station and
    window read the stored bytes and return `parse` of those, without
    fetching. `parse` therefore runs on every call, and the file holds the
    unparsed response, which any other tool can read. For a project that
    should also skip the transform on a hit, see
    [`TransformCache`][pysmo.tools.project.TransformCache].

    Any format works, as long as `fetch_raw` and `parse` agree (e.g.
    [`fetch_sac`][pysmo.tools.web.fetch_sac] with
    [`SAC.from_zip`][pysmo.classes.SAC.from_zip]). While a window stays
    cached it is replayed byte-for-byte; an entry evicted under a finite
    `max_bytes` is fetched again on next access.

    A response is stored only once `parse` has accepted it and it is
    non-empty, so an empty body (an FDSN `204` for a gap) or a response that
    does not parse is retried on the next call rather than cached as a
    permanent failure. Two threads racing the same uncached window both
    fetch.

    The call signature is
    [`SeismogramFetcher`][pysmo.tools.project.SeismogramFetcher], so an
    instance also serves as a
    [`PysmoProject`][pysmo.tools.project.PysmoProject]'s `fetch_seismogram`.
    A database written by pysmo's earlier `SqliteArchiveFetcher` is read
    as-is: the constructor is the same, and existing `.sqlite3` files need no
    migration.
    """

    path: Path = field(converter=Path)
    """Path to the SQLite file.

    Created on first use; its parent directory must already exist.
    """

    fetch_raw: RawFetcher
    """Fetches the raw response for a station and time window."""

    parse: RawParser
    """Parses a raw response into a `Seismogram`. Runs on every call."""

    wal: bool = False
    """Enable SQLite WAL mode (local disk only)."""

    max_bytes: PositiveInt | None = field(
        default=None,
        validator=validators.optional(validators.gt(0)),
    )
    """Cap on the total stored size in bytes; see
    [`BlobCache.max_bytes`][pysmo.tools.cache.BlobCache.max_bytes]. `None`
    (the default) means no cap, so a cached window is never evicted and
    re-fetched.
    """

    _cache: BlobCache = field(init=False, repr=False, eq=False)

    def __attrs_post_init__(self) -> None:
        """Build the inner store (which also checks `path`'s parent exists)."""
        self._cache = self._build_cache()

    def _build_cache(self) -> BlobCache:
        return BlobCache(
            path=self.path,
            encoding_version=_FETCH_ENCODING_VERSION,
            wal=self.wal,
            max_bytes=self.max_bytes,
        )

    def __getstate__(self) -> dict[str, Any]:
        """Drop the inner store; it is rebuilt from the plain fields on unpickling."""
        state = attrs_getstate(self, {})
        del state["_cache"]
        return state

    def __setstate__(self, state: dict[str, Any]) -> None:
        """Restore the plain fields, then rebuild the inner store."""
        attrs_setstate(self, state)
        self._cache = self._build_cache()

    def close(self) -> None:
        """Close the inner store's connection, if one is open."""
        self._cache.close()

    def __call__(
        self, station: Station, starttime: pd.Timestamp, endtime: pd.Timestamp
    ) -> Seismogram:
        """Return the `Seismogram` for `station` and window, from cache when possible.

        Args:
            station: Station to fetch data for.
            starttime: Start of the requested window (UTC).
            endtime: End of the requested window (UTC).

        Returns:
            The parsed seismogram: from the file on a hit, freshly fetched
            and stored on a miss.
        """
        key = _fetch_key(station, starttime, endtime)
        cached = self._cache.peek(key)
        if cached is not None:
            return self.parse(cached)
        raw = self.fetch_raw(station=station, starttime=starttime, endtime=endtime)
        seismogram = self.parse(raw)  # an unparseable response is not cached
        if raw:
            self._cache.put(key, raw)
        return seismogram

fetch_raw instance-attribute

fetch_raw: RawFetcher

Fetches the raw response for a station and time window.

max_bytes class-attribute instance-attribute

max_bytes: PositiveInt | None = field(
    default=None,
    validator=validators.optional(validators.gt(0)),
)

Cap on the total stored size in bytes; see BlobCache.max_bytes. None (the default) means no cap, so a cached window is never evicted and re-fetched.

parse instance-attribute

parse: RawParser

Parses a raw response into a Seismogram. Runs on every call.

path class-attribute instance-attribute

path: Path = field(converter=Path)

Path to the SQLite file.

Created on first use; its parent directory must already exist.

wal class-attribute instance-attribute

wal: bool = False

Enable SQLite WAL mode (local disk only).

__attrs_post_init__

__attrs_post_init__() -> None

Build the inner store (which also checks path's parent exists).

Source code in src/pysmo/tools/cache.py
def __attrs_post_init__(self) -> None:
    """Build the inner store (which also checks `path`'s parent exists)."""
    self._cache = self._build_cache()

__call__

__call__(
    station: Station,
    starttime: Timestamp,
    endtime: Timestamp,
) -> Seismogram

Return the Seismogram for station and window, from cache when possible.

Parameters:

Name Type Description Default
station Station

Station to fetch data for.

required
starttime Timestamp

Start of the requested window (UTC).

required
endtime Timestamp

End of the requested window (UTC).

required

Returns:

Type Description
Seismogram

The parsed seismogram: from the file on a hit, freshly fetched

Seismogram

and stored on a miss.

Source code in src/pysmo/tools/cache.py
def __call__(
    self, station: Station, starttime: pd.Timestamp, endtime: pd.Timestamp
) -> Seismogram:
    """Return the `Seismogram` for `station` and window, from cache when possible.

    Args:
        station: Station to fetch data for.
        starttime: Start of the requested window (UTC).
        endtime: End of the requested window (UTC).

    Returns:
        The parsed seismogram: from the file on a hit, freshly fetched
        and stored on a miss.
    """
    key = _fetch_key(station, starttime, endtime)
    cached = self._cache.peek(key)
    if cached is not None:
        return self.parse(cached)
    raw = self.fetch_raw(station=station, starttime=starttime, endtime=endtime)
    seismogram = self.parse(raw)  # an unparseable response is not cached
    if raw:
        self._cache.put(key, raw)
    return seismogram

__getstate__

__getstate__() -> dict[str, Any]

Drop the inner store; it is rebuilt from the plain fields on unpickling.

Source code in src/pysmo/tools/cache.py
def __getstate__(self) -> dict[str, Any]:
    """Drop the inner store; it is rebuilt from the plain fields on unpickling."""
    state = attrs_getstate(self, {})
    del state["_cache"]
    return state

__setstate__

__setstate__(state: dict[str, Any]) -> None

Restore the plain fields, then rebuild the inner store.

Source code in src/pysmo/tools/cache.py
def __setstate__(self, state: dict[str, Any]) -> None:
    """Restore the plain fields, then rebuild the inner store."""
    attrs_setstate(self, state)
    self._cache = self._build_cache()

close

close() -> None

Close the inner store's connection, if one is open.

Source code in src/pysmo/tools/cache.py
def close(self) -> None:
    """Close the inner store's connection, if one is open."""
    self._cache.close()

RawFetcher

Bases: Protocol

A callable that returns the raw bytes for a station and time window.

Called with keyword arguments only, matching pysmo's fetch functions (fetch_sac, fetch_mseed, fetch_geocsvseismogram), which can be passed straight in as FetchCache.fetch_raw.

Methods:

Name Description
__call__

Fetch raw bytes for station over starttime to endtime.

Source code in src/pysmo/tools/cache.py
class RawFetcher(Protocol):
    """A callable that returns the raw bytes for a station and time window.

    Called with keyword arguments only, matching pysmo's fetch functions
    ([`fetch_sac`][pysmo.tools.web.fetch_sac],
    [`fetch_mseed`][pysmo.tools.web.fetch_mseed],
    [`fetch_geocsvseismogram`][pysmo.tools.web.fetch_geocsvseismogram]),
    which can be passed straight in as `FetchCache.fetch_raw`.
    """

    def __call__(
        self, *, station: Station, starttime: pd.Timestamp, endtime: pd.Timestamp
    ) -> bytes:
        """Fetch raw bytes for `station` over `starttime` to `endtime`."""
        ...

__call__

__call__(
    *,
    station: Station,
    starttime: Timestamp,
    endtime: Timestamp
) -> bytes

Fetch raw bytes for station over starttime to endtime.

Source code in src/pysmo/tools/cache.py
def __call__(
    self, *, station: Station, starttime: pd.Timestamp, endtime: pd.Timestamp
) -> bytes:
    """Fetch raw bytes for `station` over `starttime` to `endtime`."""
    ...