fix: cache pyannote diarization pipeline to stop memory leak / OOM

The meeting pipeline created a fresh Diarizer per recording, each loading
the multi-GB pyannote speaker-diarization model anew (api/pipeline.py).
Whisper and Ollama run remotely in this deployment, so pyannote was the
only heavy in-process consumer. Reloading it per recording leaked CPU
memory (torch reference cycles + glibc arena fragmentation) that was
never returned to the OS, climbing to a 37 GB peak over a multi-day run
until the kernel OOM-killed the service.

Cache the loaded pipeline on the class and reuse it across Diarizer
instances, mirroring TranscriptionEngine._model. RSS now stays flat.
This commit is contained in:
2026-07-22 09:03:42 +02:00
parent 8ec9044c75
commit 6b0ee60d94
2 changed files with 36 additions and 3 deletions
+23
View File
@@ -38,3 +38,26 @@ def test_diarizer_requires_hf_token():
from diarization import Diarizer
with pytest.raises(ValueError, match="hf_token"):
Diarizer(hf_token="")
def test_pipeline_loaded_once_across_instances():
"""The heavy pyannote pipeline must be loaded once and shared, not reloaded
per recording — reloading leaks torch/CPU memory and OOM-kills the service."""
import sys, types
from diarization import Diarizer
Diarizer._shared_pipeline = None # reset shared cache for the test
fake_module = types.ModuleType("pyannote.audio")
fake_pipeline_cls = MagicMock()
fake_pipeline_cls.from_pretrained.return_value = MagicMock(name="loaded_pipeline")
fake_module.Pipeline = fake_pipeline_cls
with patch.dict(sys.modules, {"pyannote.audio": fake_module}):
first = Diarizer(hf_token="tok")._load_pipeline()
second = Diarizer(hf_token="tok")._load_pipeline()
assert first is second
fake_pipeline_cls.from_pretrained.assert_called_once()
Diarizer._shared_pipeline = None # avoid leaking mock into other tests