Skip to content
Kaamless
Book a demo

What I automate / Monitoring, alerts and on-call

Finding out before your customer does

There are two common states. The first is no monitoring: you find out something is broken when a customer tells you. The second is worse in a subtle way: so many alerts that everyone has learned to ignore them, which means the real one is also ignored.

Good monitoring for a small team is not a large platform. It is a small number of checks that mean something, routed to a person who can act, with enough context in the message that they know what to do before opening a laptop.

The test is simple and uncomfortable: when the last alert fired, did somebody act on it within minutes? If not, the alerting is decorative.

You probably know this is you if

  • You have learned about an outage from a customer
  • The alerts channel scrolls past unread
  • The same alert fires several times for one underlying problem
  • Nobody is sure who is supposed to respond outside office hours
  • Writing up what happened after an incident is done from memory

The jobs in this area

  • Finding out from a customer that it is down

    Health checks from outside your network, alerting the right person on the right channel within a minute.

  • So many alerts that nobody reads them

    Alerts grouped, de-duplicated and routed by severity, so the ones that wake someone up are the ones that should.

  • Writing the incident timeline from memory

    The timeline assembled from what actually happened, ready to edit rather than reconstruct.

  • Answering 'is it just me?' in chat

    A status page that updates from the same checks, so the question answers itself.

  • The same fix applied manually every time

    A known, safe remedy triggered automatically, with a record of when it ran and whether it worked.

How a build like this goes

  1. 1

    Start from outside your network, checking what a customer actually experiences, because internal checks can be green while the service is unreachable.

  2. 2

    Cut the noise before adding anything: group related alerts, de-duplicate, and delete the ones nobody has acted on in six months. An alert nobody acts on is a bug in the alerting.

  3. 3

    Route by severity, to a channel people watch, with the context needed to act in the message itself.

  4. 4

    Then the extras that pay off later — a status page fed by the same checks, and an incident timeline assembled from what happened rather than remembered.

A worked example

External checks for a customer-facing service

Before
Two outages last year found by customers, each running about 90 minutes before anyone noticed.
After
Checks every minute from outside, alerting the person on call. Detection drops from 90 minutes to under two.

This one resists neat arithmetic and I would rather say so than invent a number. The build is quick-win band. Whether it is worth it depends on what 90 minutes of silent downtime costs you — in revenue, in support calls, or in a customer deciding you are unreliable. Most people can answer that question in one sentence.

Those are illustrative figures, not a quote and not a promise. Your own numbers are the only ones that matter — the calculator does the same arithmetic on them.

What this normally costs

A job in this area is usually a quick win: ₹18,000 to ₹45,000, ready in 3 to 7 working days. The exact number comes from the audit, in writing, before anything is built.

The audit itself is ₹9,999 and comes off the first build in full. If it turns out this is not worth automating for you, that is what the report will say.

See the full price list

Questions about this kind of work

We already have a monitoring tool and nobody uses it.

Very common, and usually a routing and noise problem rather than a tooling one. Often the fix is to keep the tool and rebuild what it alerts on, which is cheaper than replacing it.

Do we need to pay for a monitoring service?

Sometimes, and the free tiers are genuinely sufficient for a small number of checks. I will tell you where a paid tier earns its cost and where it does not.

Can this page our phones at night?

It can. Whether it should is a management decision, not a technical one — and if nobody is actually going to act at 3am, honest silence is better than an alert everyone ignores.

Tell me about your version of this

Anything on this list, or something that is not on it. Ask, and you will get a straight answer about whether it can be built.

What happens today, who does it, and how often. A few lines is plenty.

I reply within one working day. No mailing list, no follow-up you did not ask for.

Tell me the job you are tired of

One call, half an hour. You will get a straight answer about whether it is worth automating — including 'no' if that is the answer.