fix: cache pyannote diarization pipeline to stop memory leak / OOM

The meeting pipeline created a fresh Diarizer per recording, each loading
the multi-GB pyannote speaker-diarization model anew (api/pipeline.py).
Whisper and Ollama run remotely in this deployment, so pyannote was the
only heavy in-process consumer. Reloading it per recording leaked CPU
memory (torch reference cycles + glibc arena fragmentation) that was
never returned to the OS, climbing to a 37 GB peak over a multi-day run
until the kernel OOM-killed the service.

Cache the loaded pipeline on the class and reuse it across Diarizer
instances, mirroring TranscriptionEngine._model. RSS now stays flat.
This commit is contained in:
2026-07-22 09:03:42 +02:00
parent 8ec9044c75
commit 6b0ee60d94
2 changed files with 36 additions and 3 deletions
+13 -3
View File
@@ -2,19 +2,29 @@ import asyncio
class Diarizer:
# The pyannote pipeline holds multi-GB torch models. A fresh instance is
# created per recording (see api/pipeline.py), so cache the loaded pipeline
# on the class and reuse it — reloading per recording leaks CPU memory
# (torch reference cycles + glibc arena fragmentation) and eventually OOMs
# the long-running service. Mirrors TranscriptionEngine._model.
_shared_pipeline = None
def __init__(self, hf_token: str):
if not hf_token:
raise ValueError("hf_token is required for diarization")
self._hf_token = hf_token
self._pipeline = None
self._pipeline = None # per-instance override (used by tests)
def _load_pipeline(self):
if self._pipeline is None:
if self._pipeline is not None:
return self._pipeline
if Diarizer._shared_pipeline is None:
from pyannote.audio import Pipeline
self._pipeline = Pipeline.from_pretrained(
Diarizer._shared_pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-3.1",
token=self._hf_token,
)
self._pipeline = Diarizer._shared_pipeline
return self._pipeline
async def diarize(self, wav_path: str) -> list[tuple[float, float, str]]: