• Home
  • Blog
  • Projects
  • Bookshelf
  • About
← Back to all posts

Service Monitors and Observability

Published ....
Last modified ....

Share this post on BlueskySee discussion on Bluesky


Thanks to Scott Kaye for reviewing an early draft of this post

Awhile back I tweeted made this post on X about being intentional when creating service monitors:

Twitter

This became my go-to question when we were defining additional monitors for our services over the past month or so:



“If you get paged by this at 2am, do you know what it means/what to do?”



It’s a really solid litmus test to ensure monitors are clear, concise, and actionable https://t.co/YpFhbCW34G

— Matt Hamlin (@immatthamlin)

July 29, 2023

I figured I would expand a bit on it within a blog post since I have thought a decent amount about it since I posted it.

For those that don't know, service monitors are automatic "tests" that can be used to determine the health of a system in production. They can be configured for just about anything, latency and up-time are generally the most popular monitors.

Traditionally, monitors are configured to automatically raise an incident, which usually pages someone that is currently on-call for the service.

However, as uncle Ben says - with great power comes great responsibility.

It can be nerveracking to be on-call for large and complex services, especially if you're new to the team that owns such a service. Now compound that with being paged due to a monitor that is missing context about what's actually wrong with the service and what to do to resolve it.

This is what I was speaking to in the above post on Twitter/X, at work we rapidly spun up a ton of new monitors for our core service after a particularly interesting series of incidents (maybe I'll write about those in the near future). However, at the start we weren't necessarily thinking about them from the framing of the original post and instead we were mainly thinking about adding more observability to the systems and service as a whole to help cover some of the gaps identified in the previous incidents.

Fortunately, wiser minds on the team prevailed, and we started to become more critical about the monitors we were creating. We started to add more context to the monitors, outlining what it means for a particular monitor to be tripped, and what one should do to help resolve the issue if it is happening.

We even started to cull back some of the many monitors we created for the service, this may seem a bit counterintuitive but another risk of creating monitors is adding noise to the engineers on-call. We found that we'd be paged for issues that would auto-resolve in a short amount of time, or even those that we couldn't even do anything about during the time of the incident.

All of these patterns made it more difficult to support our services rather than making it easier.

As with everything, there's nuance. Service monitors offer a lot of benefits and help improve overall service and system observability. However they must be applied appropriately. Try to remember - what can this monitor tell me when I get paged at 2 or 3 in the morning?



Tags:

Web DevelopmentmicropostObservability

Loading...

Related Posts

Web Development

Website Redesign v11

....

I've once again updated my personal website, this time powered by my own RSC framework and a custom CMS solution as well!

Onboarding Your New AI Teammate

....

Are we reinventing the wheel when it comes to onboarding new AI agents to a codebase, when we already have primitives available for onboarding humans?

Quick Tip - Theme Aware Images

....

Have you ever found the need to change the image you render on a web page based on the current preferred color scheme of your theme?

Async Class Creation In JavaScript

....

Have you ever wanted to create a class in JavaScript or TypeScript but also have the initialization be async? Here's a quick tip on a pattern that I've used in the past!

Website Redesign v10

....

I recently launched a rewrite and redesign of this personal website, I figured I'd talk a bit about the changes and new features that I added along the way!

Server Side Rendering Compatible CSS Theming

....

A quick tip to implementing CSS theming that's compatible with Server Side Rendered applications!

Podcasting By Hand

....

A brief overview on how we launched The Bikeshed Podcast, including a deep dive in our recording and distribution workflows!

Quick Tip - Specific Local Module Declarations

....

A quick tip outlining how to provide specific TypeScript type definitions for a local module!

On File-System Routing Conventions

....

Some rough thoughts on building a file-system routing based web application

You're Building Software Wrong

....

Slicing software: why vertical is better than horizontal.

Single File Web Apps

....

What if you could author an entire web application in a single file?

Resetting Controlled Components in Forms

....

A quick way to handle resetting internal state in components when a parent form is submitted!

A Quick Look at Import Maps

....

A brief look at Import Maps and package.json#imports to support isomorphic JavaScript applications!

Recommended Tech Talks

....

A collection of tech talks that I regularly re-watch and also recommend to everyone!

Request for a (minimal) RSC Framework

....

Some features and functionality that I'd like within a React Server Component compatible framework.

Bluesky Tips and Tools

....

A (running) collection of Bluesky tips, tools, packages, and other misc things!

The Bookkeeping Pattern

....

A quick look at a small but powerful pattern I've been leveraging as of late!

TSLite

....

A proposal for a minimal variant of TypeScript!

Monorepo Tips and Tricks

....

Sharing a few core recommendations when working within monorepos to make your life easier!

Next.js with Deno v2

....

This is a quick post noting that Next.js should now work with Deno v2!

Don't Break the Implicit Prop Contract

....

React components have a fundamental contract that is often unstated in their implementation, and you should know about it!

A Better useSSR Implementation

....

Replace that old useState and useEffect combo for a new and better option!

My Current Dev Setup

....

A quick look at the applications and tools that I (generally) use day to day for web development!

Abstract Your API

....

Proposing a solution for sharing core "business" logic across services!

Tip: Request and Response Headers

....

There's a common gotcha when creating Web Request and Response instances with Headers!

Using Feature Toggles to De-risk Refactors

....

Feature toggles are often underused by most software development teams, and yet offer so much value during not only feature development but also refactors

Hohoro

....

A quick introduction to my new side project, hohoro. An incremental JS/TS library build tool!

Custom Favicon Recipes

....

Two neat tricks for enhancing your site's favicon!

Corporate Sponsored OSS

....

The various risks and pitfalls of open source software run by corporations.

The Library-Docs Monorepo Template

....

A monorepo template for managing a library and documentation together.

Building Better Beacon

....

How we solved an almost show-stopping production bug, and how you can avoid it in your own projects.

Project Deep Dive: Tails

....

A(nother) deep dive into one of my recent side projects; tails - a plain and simple cocktail recipe app.

Churn Anxiety

....

When did semver major changes become so scary?

On Adopting CSS-in-JS

....

A brief recap of how Wayfair changed it's CSS approach not once but twice in the span of 5 years!

Project Deep Dive: Microfibre

....

A deep dive into one of my recent side projects; microfibre - a minimal text posting application

Pair Programming

....

Pair programming can be good sometimes - but not all the time

Suspense Plus GraphQL

....

A few thoughts on using Suspense with GraphQL to optimize application data loading

You've Launched, Now What?

....

A few thoughts on what to do after you launch a new project

Taking a Break

....

A few quick thoughts on burn out and taking a break

Managing Complex UI Component State

....

A few thoughts on managing complex UI component state within React

Understanding React 16.3 Updates

....

A quick overview of the new lifecycle methods introduced in React 16.3

CSS in JS

....

A few thoughts and patterns for using styled-jsx or other CSS-in-JS solutions

Redesign v6

....

A few thoughts on the redesign of my personal site, adopting Next.js and deploying via Now

JavaScript Weirdness

....

A few weird things about JavaScript

Calendar

....

Building a calendar web application

micropost

Quick-Tip: Codex Chief of Staff Thread

....

Tip: Secrets for Cloudflare Workers

....

Tip: GitHub Created Date Filtering

....

Link: Next.js Is Infuriating

....

Dependabot Hell

....

Link: What the hell is going on right now?

....

Recipe: Horchata Protein Latte

....

Zombie Retros

....

Link: How I build software quickly

....

Vacation (and streaks)

....

Polish is Important

....

<Blank> Driven Development

....

All Documentation Should Be Dated

....

Adding Microposts

....

Being Unopinionated

....

It's fine for a library to express some opinions about how it should be adopted and how the overall workflow/application in which it is adopted should function. However, it's false advertising to say that it is unopinionated.

Stop Snacking

....

No I don't mean those Milano cookies you keep taking from the office snack wall either (although you should probably stop snacking on those as often as well).

No Process is Invisible Process

....

Low/no process workflow wasn't actually no process, it was only an "invisible" process. An implicit contract with everyone on the team to do that async workflow on their own time.