An LLM pipeline suddenly starts producing strange responses, even though accuracy doesn't decrease and latency doesn't increase. Standard metrics remain silent, and the responses become increasingly nonsensical. Often, the culprit is "silent" token drift: a subtle shift in the distribution of embeddings without ground truth. The causes can be: updating the model, changing the tokenizer, or the input data gradually shifting.
➜ Spectral analysis of embeddings using FFT
One effective approach is spectral analysis without labeling, relying solely on statistics. I collect embeddings from production logs (each request is a vector with 768 or 1024 dimensions). I reduce the dimensionality to 50-100 components using PCA – not for visualization, but to remove noise and retain the essential information. For each component, I calculate the FFT, obtaining a power spectrum. I monitor shifts in peak frequencies: for example, I look at the ratio of energy in low frequencies (0.1-0.3 Hz) to high frequencies. In a stable state, the spectrum maintains a pattern: 60% of the energy is in the low frequencies. After an update, the peak shifts by 0.5 Hz – this is a signal.
import numpy as np
from scipy.fft import fft
def detect_drift(embeddings_batch, baseline_spectrum, threshold=0.15):
avg_embed = np.mean(embeddings_batch, axis=0)
spectrum = np.abs(fft(avg_embed))[:len(avg_embed)//2]
spectrum = spectrum / np.sum(spectrum)
diff = np.abs(spectrum - baseline_spectrum)
drift_score = np.max(diff)
return drift_score > threshold
➜ Setting the threshold and a common mistake
I empirically determine the threshold based on the 95th percentile of historical data from the past week or month. The method doesn't require labeling, only embedding logs. However, this is not a silver bullet: if the drift is gradual, the threshold will need to be recalculated. A common mistake is to calculate the FFT on the full dimensionality of the embeddings without PCA. Noise overwhelms the signal, and the threshold becomes useless. A practical tip: when the system triggers, check the tokenizer – compare tokenizer.encode("test") before and after the update, or look at the distribution of out-of-vocabulary (OOV) tokens. Often, the cause is a change in the frequency of rare tokens: Unicode characters, special symbols, emojis that the model rarely saw before.
➜ Trade-offs and engineering limitations
The method is computationally inexpensive (O(n log n) per batch) and suitable for real-time monitoring, but it doesn't provide an answer as to "what exactly is broken." It's an indicator for MLOps: it signals "go check" before users start reporting issues. The main trade-off is sensitivity to the batch size: a batch that is too small (less than 10 requests) can lead to false positives due to noise, while a batch that is too large (more than 1000) can delay detection by hours. For production, I choose a batch size of 50-100 embeddings and check the system every minute.
☞ Conclusion: Spectral analysis of embeddings using FFT with PCA reduction is a cheap and label-free method for detecting "silent" token drift in LLM pipelines, which is crucial before users start complaining.
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
