{"id":5267,"date":"2026-10-08T19:08:34","date_gmt":"2026-10-08T19:08:34","guid":{"rendered":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/08\/attaching-evidence-to-alerts-an-enrichment-sidecar-for-alertmanager\/"},"modified":"2026-10-08T19:08:34","modified_gmt":"2026-10-08T19:08:34","slug":"attaching-evidence-to-alerts-an-enrichment-sidecar-for-alertmanager","status":"publish","type":"post","link":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/08\/attaching-evidence-to-alerts-an-enrichment-sidecar-for-alertmanager\/","title":{"rendered":"Attaching Evidence to Alerts: an Enrichment Sidecar for Alertmanager"},"content":{"rendered":"<div><img data-opt-id=1639743794  fetchpriority=\"high\" decoding=\"async\" width=\"770\" height=\"330\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/Monitoring-Alert-Workflow-Dashboard-Large-e1791482957391.jpeg\" class=\"attachment-large size-large wp-post-image\" alt=\"\" \/><\/div>\n<p><img data-opt-id=873560358  fetchpriority=\"high\" decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/Monitoring-Alert-Workflow-Dashboard-Large-150x150.jpeg\" class=\"attachment-thumbnail size-thumbnail wp-post-image\" alt=\"\" \/><\/p>\n<p data-pm-slice=\"1 1 []\">Every on-call engineer knows the ritual. A page lands in Slack: the 5xx rate on a service is above the threshold. The message tells you that something is wrong and nothing about what. So you open Kibana or Loki in another tab, set the time window to the last few minutes, filter by the service, and start reading. In my team this took two or three minutes on every alert, and these are the worst minutes of the incident, because nothing is being diagnosed yet.<\/p>\n<p>The obvious request is to put the last few error lines into the alert message itself. I tried to do this inside Alertmanager, and it cannot be done there. It helps to understand why before building anything around it.<\/p>\n<h3>Why Alertmanager Cannot Do It<\/h3>\n<p>Alertmanager renders notifications through Go templates, and a template can only use what is already attached to the alert: the labels and the annotations which Prometheus put there when the rule fired. There is no mechanism which would call Elasticsearch, Loki or any other HTTP API at the moment a notification is sent. Annotations are rendered from the alerting rule at evaluation time, inside Prometheus, and Prometheus does not talk to log stores either. Fetching data at send time is simply not a capability of the design.<\/p>\n<p>So the logs have to come from a separate process which stands next to Alertmanager.<\/p>\n<h3>The Sidecar<\/h3>\n<p>What runs in production now is a small service next to Alertmanager, and its shape came out of two failed attempts, so I will describe it through them.<\/p>\n<p>My first attempt was a bot which received alerts and posted them to Slack itself, with the logs already attached. It looked good in a demo, but it was a mistake because the bot had become the paging path. Any bug in my code, any timeout of the log store, and nobody gets paged. I threw that version away quickly. Alertmanager has to stay the thing which posts the alert, exactly as before, and whatever adds context has to be allowed to die without anyone noticing. So the enricher hangs off a second receiver: the route which sends the alert to Slack also duplicates it to a webhook, and behind the webhook stands my service. When it is down, the page arrives as it always did, just without the extra.<\/p>\n<p>The second attempt posted the logs as its own message into the channel, right after the alert. Within a day the channel was hard to read: two messages per alert, and during an alert storm they interleaved, so the logs stood three messages away from their page. The fix was to reply into the thread of the message which Alertmanager had already posted. The channel stays as it was, one message per alert, and whoever opens the incident an hour later finds the logs from the exact minute of the page sitting under it.<\/p>\n<h3>Finding the Message to Reply To<\/h3>\n<p>That thread reply is the one interesting engineering problem in the whole service. The enricher receives the alert from Alertmanager, and it has to find the Slack message which corresponds to this alert, and nobody tells it which one that is. When the webhook fires, the service reads the last couple of minutes of channel history and looks for a message which carries the same alert fingerprint. Our Slack template rendered the fingerprint into the message, so the match was exact; before I added that, I matched on the label set which appears in the message text, and that worked too. Once the message is found, the reply is one API call with thread_ts.<\/p>\n<p>Then production corrected the details. Alertmanager re-sends a firing alert on repeat_interval, and my version at first posted the same log dump under the same message again and again, so now the service remembers which fingerprints it has already enriched. The lookback for matching is limited to about two minutes, since replying to some old message which happens to carry the same labels is worse than staying silent. The query into the log store is built from the labels of the alert (service, namespace, severity) and takes the last few error-level lines, ten or so. I keep it that short on purpose: the reply is a hint for the person, and the full picture is in the log store anyway. One more thing I did not expect: an empty result is also worth posting. When the reply says there were no matching error logs in the last five minutes, the person on call can broaden the investigation to the infrastructure or the alerting rule itself, and this saves the same minutes.<\/p>\n<h3>What It Changed<\/h3>\n<p>The service is a few hundred lines, and in operation it has been boring, and boring is what I want from anything standing near paging. The visible change for the team: when you open Slack on a page, the first evidence is already under it, and the phase of \u201cwhich dashboard, which index, which time window\u201d is gone from incidents. Time to diagnosis dropped on every incident which had logs behind it, without any change to the services themselves.<\/p>\n<p>Logs turned out to be only the first thing worth attaching. The same service can reply with anything which Alertmanager cannot know at render time: a runbook link picked by label, the marker of the latest deploy of the affected service (the first question in half of the incidents anyway), or a dashboard link already scoped to the incident window. Alertmanager alerts, and the sidecar carries what the responder reaches for in the first minute.<\/p>\n<p>If you run Prometheus with Slack, the whole thing is one small service plus one extra webhook receiver in the Alertmanager config. Do the idempotency bookkeeping first; it is the first thing which bites.<\/p>\n<p><em>Daniil Romashov is a Site Reliability \/ DevOps engineer working on reliability and observability for high-load systems. Open-source work: github.com\/youngpabl0.<\/em><\/p>\n<p><a href=\"https:\/\/devops.com\/attaching-evidence-to-alerts-an-enrichment-sidecar-for-alertmanager\/\" target=\"_blank\" class=\"feedzy-rss-link-icon\">Read More<\/a><\/p>\n<p>\u200b<\/p>","protected":false},"excerpt":{"rendered":"<p>Every on-call engineer knows the ritual. A page lands in Slack: the 5xx rate on a service is above the [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":5268,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[5],"tags":[],"class_list":["post-5267","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops"],"_links":{"self":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5267","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/comments?post=5267"}],"version-history":[{"count":0,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5267\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media\/5268"}],"wp:attachment":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media?parent=5267"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/categories?post=5267"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/tags?post=5267"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}