Alerting Rules: "Something Broke" vs. "Users Are Affected"
Cause-based alerts (high CPU, full disk, failed job) wear out on-call engineers without answering whether users are actually affected. How to move to SLO-based burn rate alerting, and the traps along the way.
In any system you can build two kinds of alerts. The first says "this component is behaving abnormally": CPU crossed ninety percent, disk filled up, the queue backlog grew, a cron job exited non-zero. The second says "this group of users is having a bad time right now": the request failure rate crossed a defined threshold, the latency distribution moved outside an acceptable range. Most teams pile up alerts in the first category because it is easy to set up and the metrics are already there. The result is a familiar story: the on-call phone rings at 3am over a disk alert, half an hour gets spent, and it turns out disk usage climbed to eighty five percent while not a single user request failed. In the morning nobody can answer whether anything actually happened last night.
The two approaches are not interchangeable, but mixing them up makes both useless. Cause-based alerting, used correctly, gives early warning: you see disk growth before it fills up, you catch a memory leak before it reaches production traffic. Used incorrectly it produces alert fatigue. A team that gets fifty notifications a day stops reading by the forty ninth and loses the ability to tell which one is real. This is the consistent finding across large-scale on-call reports published in recent years: as notification volume rises, average response time does not shrink, it grows.
Making user impact measurable
The precondition for outcome-based alerting is turning the phrase "users are affected" into a number. Three terms from site reliability engineering practice do the job: the Service Level Indicator (SLI), its target value (SLO), and the failure budget the target allows (error budget). Typical SLIs for a web service are the fraction of successful requests, the percentage of requests completing under a given latency threshold, and the rate at which a queued job gets processed within a defined time window. The SLO is the target that indicator must hold over a rolling window, say a thirty day period, for example 99.9 percent success. The remaining 0.1 percent is the error budget the team consciously accepts, and it gets spent on decisions like taking deploy risk or opening a maintenance window.
The critical distinction is this: CPU at ninety percent is not an SLI, because its direct relationship to user experience has never been established. A system can run at ninety percent CPU while every request still returns correctly and on time. An SLI is what the user directly feels: did the request succeed, how long did it take, was the data correct.
Burn rate alerting: when and how severe
Turning an SLO into a single-threshold alert ("warn if monthly error rate exceeds 0.1 percent") does not work in practice, because you would have to wait until the end of the month, and you would miss a fast degradation entirely. The technique used instead is multi-window, multi-burn-rate alerting. The idea is simple: measure how fast your error budget is being consumed. The budget is designed to be spent over thirty days under normal conditions; if the current error rate would exhaust it within one hour, that is a burn rate forty four times the normal pace and demands immediate action. The same calculation is repeated over six hour and one day windows to catch slower, sustained degradation as well.
A sample rule in Prometheus looks like this:
groups:
- name: slo-burn-rate
rules:
- alert: ErrorBudgetBurnFast
expr: |
(
sum(rate(http_requests_total{code=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) > (14.4 * 0.001)
for: 2m
labels:
severity: page
annotations:
summary: "Error budget burning fast, 2% of the monthly budget spent in the last hour"
Here 0.001 represents the monthly target error rate (a 99.9 percent SLO), and the 14.4 multiplier represents the burn coefficient at which the entire budget would be exhausted in roughly two days at this rate. The short window (5 minutes) is checked alongside the long window (1 hour) so a brief spike does not trigger a false page, while a genuine degradation still cannot hide for a full hour unnoticed. Slower burn rates (6x, 3x, 1x) follow the same pattern with longer windows and are typically routed to ticket-level notification rather than paging.
Do not throw away cause-based alerting
A common mistake once teams move to outcome-based alerting is dropping cause-based monitoring entirely. The two serve different purposes. An outcome-based alert tells you "act now" and carries the authority to wake someone up. Cause-based metrics are used after the page fires, to find the root cause: disk usage, queue depth, connection pool saturation, retry counts. The correct architecture routes cause-based signals to dashboards and low-priority notifications, and reserves paging for signals that prove actual user impact. In Alertmanager this separation is done with a severity label: rules labeled severity=page route to a receiver that reaches the on-call engineer directly, while severity=ticket rules route to the issue tracker.
Traps
The first trap is the "warn and continue" pattern that swallows a failure and keeps going. When a step fails but the rest of the pipeline keeps running, and the only symptom is a log line, that log line gets read by nobody and the dashboard stays green. At every such point two questions need answers: is continuing while this step is broken actually better than not running at all, and if you continue, can you measure what data the next step is now operating on? If both answers are no, the pipeline should stop and make noise.
The second trap is never verifying that the alert actually fires. Computing an SLI from the wrong metric source, querying it with the wrong label set, or reading data written into a different namespace than the one being queried, silently turns the alert into something that will never trigger. The only reliable test is deliberately producing a real degradation: inject errors into test traffic and confirm the alert actually fires, routes to the right channel, and triggers at the correct threshold. This is proving the monitoring chain works end to end by causing the failure in an isolated setting; it is not a one-time setup step, it is something to repeat every time the alerting logic changes.
The third trap is picking windows too narrow or thresholds too loose. A window that is too narrow triggers on ordinary traffic spikes, and the on-call engineer learns to ignore it; a threshold too loose lets a real outage go unnoticed for hours. Window and multiplier choices should be backtested against historical traffic: look at when the new rule would have fired during past real incidents, and eliminate rules that fire either too late or too early.
When not to adopt this approach
This investment is not necessary for every system. An early-stage product with low traffic and few users does not have enough sample size to compute a meaningful SLI; a 99.9 percent target on a service handling a few hundred requests a day cannot be statistically distinguished from noise. In that situation simple threshold alerts (is the service up, do core endpoints respond) are sufficient, and the return on building SLO infrastructure does not justify the engineering time spent. As traffic and user count grow, especially once multiple teams own different parts of the same system, impact-based alert prioritization starts to show its real value.
[OK] 25 records · 8 featured · 1 retainer · 16 standard
/02 services
services.tree
> note
· Prices are in USD. Volume discounts of 7% for 6-12 day projects and 14% for 13+ day projects apply. Programs of 100+ person-days are priced individually. Prices are negotiable based on scope, urgency and long-term collaboration.
> ls ~/products/ --open-source
[OK] 2 entries · 1 live · 1 pre-release
/03 products
products.list
liveopen source · mit/prod/01
filex
the self-hosted file manager that embeds anywhere
single go binary · vue/react/web-component embed · 5 storage drivers · realtime collab · rbac · native multi-tenancy · mcp server for ai agents
We build infrastructure, automate everything, and keep systems running.
BRF Tech is a Bursa-based DevSecOps and software consulting firm. We operate across 25 service areas, from Kubernetes cluster management to serverless platform setup, CI/CD pipeline design to AI agent development.
What we do is simple: we set up your systems, automate them, and make sure they won't wake you up at 3 AM. We codify your infrastructure with Terraform, move your deployments to GitOps with ArgoCD, and monitor everything with Prometheus. If something breaks — we intervene before it does.
We're against vendor lock-in. We work with open-source tools, self-hosted solutions, and industry-standard technologies. We build your own serverless platform on Knative, isolate with Kata Containers, manage your secrets with Vault. Every project is delivered with clear scope, clear timeline, clear pricing.
DevSecOps & CI/CD
Kubernetes & Serverless
Infrastructure as Code
AI & Agent (MCP/ACP)
Security & Zero Trust
Full-Stack Development
Self-Hosted Solutions
> grep -i question ~/faq.md
[OK] 3 entries · click to expand
[?]What services are included in BRF Tech's DevOps solutions?▾
We offer CI/CD pipeline setup, Kubernetes cluster management, container migration, infrastructure automation with Terraform, monitoring & alerting, in-house technology installations and DevOps consulting services.
[?]What are your software development services?▾
We offer MVP development, existing software performance optimization, backend API development (Node.js, Go, Python) and full-stack application development. Every project is delivered with minimal technical debt.
[?]Which DevOps tools do you work with?▾
From Kubernetes and Docker to Talos Linux, from Terraform and OpenTofu to Ansible, and across Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, Flux, Helm and Rancher, we work with industry-standard tooling. For observability and error tracking we use Prometheus, Grafana, Loki, Tempo, OpenTelemetry and GlitchTip; for networking, Cilium and Envoy; for test automation, Playwright, Cypress, k6, Locust, Testcontainers and Trivy. We determine the most suitable toolset together, based on your project’s needs.
Automated build, test and deployment pipelines with GitHub Actions, GitLab CI/CD.
Automated build, test and deploy pipeline setup for existing or new repositories. GitHub Actions, GitLab CI/CD or preferred tool is used. Separate workflows are defined for staging + production environments.
Container architecture migration, cluster setup and orchestration.
Production-ready Kubernetes cluster setup on bare-metal or cloud. Includes Helm charts, Ingress, TLS certificate management, namespace isolation and RBAC.
Migration of existing applications to container architecture.
Migration of existing applications to container architecture. Docker image design, Compose configuration and conversion to Kubernetes manifests. Kata Containers / gVisor can be included for advanced isolation.
In-house tools like GitLab, Nextcloud, VPN, mail server.
Installation of tools like GitLab, Mattermost, Nextcloud, mail server, VPN, internal monitoring on company servers. Deployed on Docker Compose or Kubernetes.
DB + cache + code bottleneck identification and resolution.
Profiling of existing application, identification and resolution of bottlenecks. DB query optimization, caching layer (Redis), service-level improvements.
MCP/ACP agents, RAG pipelines, multi-agent orchestration, model serving.
LLM integration (OpenAI, Claude, Gemini, local models), tool-augmented AI agents with MCP (Model Context Protocol) and ACP (Agent Communication Protocol). RAG pipeline, multi-agent orchestration, model serving (vLLM/Ollama). AI layer for existing business processes.
Event-driven data pipeline with Apache Kafka, RabbitMQ, NATS. ETL/ELT orchestration with Apache Airflow, real-time data streaming, CDC (Change Data Capture). Data lake/warehouse design, schema registry, dead letter queue management.
Packages
Kafka/RabbitMQ setup + basic producer/consumer
$3,500
5–7 person-day
ETL pipeline (Airflow + source → warehouse)
$5,200
8–12 person-day
Full event-driven architecture (CDC + streaming + DLQ)
Hands-on technical training for your teams. Docker & container fundamentals, Kubernetes operations, Terraform IaC, CI/CD best practices, AI/LLM integration workshops. Practical exercises in live lab environments, content customized by skill level.
E2E test automation with Playwright, Cypress, load testing with k6/Gatling. Pre-deploy quality gate integrated into CI/CD pipeline, test coverage reporting, visual regression testing. Reduce test writing time by up to 60% with AI-assisted test generation.
Your Collected Personal Data, Collection Method and Legal Basis
Any information that identifies or makes you identifiable is considered "personal data." When you visit our website or use our services, contact information such as your name, surname, email address, phone number, as well as your IP address and browser cookie data may be collected through automatic or semi-automatic means. This data is processed based on the legal grounds specified in Article 5 of the Personal Data Protection Law No. 6698: "being directly related to the establishment or performance of a contract" and "being mandatory for the legitimate interests of the data controller, provided that it does not harm the fundamental rights and freedoms of the data subject."
Purpose of Processing Your Personal Data
Your collected personal data is processed for the purposes of providing and improving our services, fulfilling customer requests, meeting legal obligations, conducting information security processes and managing communication activities. Your data is processed in a limited and proportionate manner for the stated purposes; when the purpose ceases to exist, data is deleted, destroyed or anonymized.
To Whom and For What Purposes Collected Personal Data May Be Transferred
Your personal data may be transferred to public institutions and organizations as required by legal regulations, to our business partners and technical infrastructure providers for the purpose of delivering services, and to lawyers and consultants in legal disputes. Transfers are carried out in accordance with Articles 8 and 9 of the Law, with necessary technical and administrative measures in place.
Your Rights as a Data Subject
In accordance with Article 11 of Law No. 6698, you have the right to: learn whether your personal data is being processed; request information if it has been processed; learn the purpose of processing and whether it is used in accordance with its purpose; know the third parties to whom it is transferred domestically or abroad; request correction if it has been processed incompletely or incorrectly; request deletion or destruction within the framework of conditions set out in Article 7 of the Law; request that the operations carried out be notified to third parties to whom data has been transferred; object to any adverse result arising from the analysis of data exclusively through automated systems; and claim compensation for damages in case of unlawful processing. You may contact us through the communication channels on our website to exercise these rights.
Terms & Conditions
Terms & Conditions
1. Service Scope and Changes
BRF Tech conducts its software development, DevOps consulting, infrastructure setup and technical support services within the framework of these terms and conditions. The scope of service is determined separately for each project and finalized through mutual agreement. BRF Tech reserves the right to update the scope, pricing and technical details of its services, provided that prior notice is given.
2. Privacy Policy
Personal data collected within the scope of our services is processed in accordance with our Privacy Policy. For detailed information, please review our Privacy Policy page.
3. Service Usage and Responsibilities
Clients agree to use the provided services solely within legal and ethical boundaries. BRF Tech reserves the right to suspend or terminate services in the event of misuse, violation of third-party rights, or use for illegal activities. The client is responsible for the accuracy and currency of the information provided within the project scope.
4. Payment Terms
Payment terms are determined on a project basis and finalized through mutual agreement at the start of the project. Unless otherwise specified, invoices are payable within 15 days of project delivery. A monthly late payment interest of 2% may apply to overdue payments.
5. Cancellation and Refund Policy
Cancellation requests must be submitted in writing at least 7 days before the service start date. For projects already in progress, billing is based on the proportion of completed work. No refunds are issued for software and infrastructure components specifically developed within the project scope.