Growing pains: An update on the ListenBrainz service status

TL;DR: The growth of ListenBrainz has caught up with us, and our limited team is working on replacing central parts of our infrastructure. In the meantime, many features are unstable.

We are victims of our success. While in the long term this is a good problem to have, the sharp increase in users over the past year and a half has left our infrastructure cracking at the seams.

The good news: Your listens are being stored and imported without issue, even if there are delays. The core service is not compromised and we ask that you please keep submitting your listens. Your stats and playlists will be back!

However, we know that every other week your statistics, weekly playlists, and other features fail to generate for everyone, and cause crashes. We know it is frustrating, and we share that feeling.

While we are aware of the issues, fixing them is far from simple and requires us to completely rework all the crucial parts of our infrastructure.

In the interest of transparency, here are our main issues and what we are doing to fix them:

How many listens?

Our database dumps have become too big for us to process.
We went from 0 to 1 billion listens in 7 years – then to 2.5 billion in the next one and a half years.

Generating and copying our (now huge) database dumps causes crashes as our limited servers run out of memory and disk space.

We are working on improvements, but for each change we need to wait two days to be able to run tests.

In addition, we are moving away from TimescaleDB, a Postgres database extension for time-series data, in favour of vanilla Postgres with a new partition scheme.

We found that TimescaleDB was not adapted for our use case of working with historical imports or deleting listens and users (and all their listens), as well as large gaps between listens, all causing some very slow queries.

Where are my stats, goddammit?

Our statistics and playlist calculation infrastructure, a Spark cluster of 5 servers, is running out of memory and crashing, from one task or another. This used to happen once every few months but is now a weekly occurrence.

This is the issue which is breaking stats, playlists, user similarity, unlinked listens, fresh releases and more.

We are moving to using Clickhouse instead for all statistics calculations, which will free up the Spark cluster to be used for generating playlists and other tasks.
This move is taking some time, as a single person in the team carries the responsibility of rewriting essential code and testing everything carefully.

This will eventually open the door to requesting stats for an arbitrary time range instead of being limited to this and last week/month/year, a hotly requested feature that is not possible with our current system.

The scraping situation

To make matters worse, the entire internet is being bombarded by unscrupulous bad actors (looking at you, AI companies) that don’t follow the rules and try very, very hard to evade any measure meant to limit them.

They rent botnets of millions of residential IPs so they can scrape our APIs and websites incessantly, over and over again, while evading detection, all for data that they could download for free.

They cause surges of 5x the usual traffic across all our projects and slow everything down on our resource-constrained infrastructure.
It also forces us to spend time dealing with these DDOS-like surges instead of working on our other pressing issues.

For ListenBrainz specifically, we have had to disable some features/endpoints that during scraping waves made the website completely unreachable for everybody.

But wait, there’s more…

The loss of our founder in late February was big blow to our team.

Rob was one of the custodians of Listenbrainz infrastructure, but also a central ListenBrainz team member.

We have had to cross-train our ListenBrainz dev team -it is only three of us- to deal with infrastructure and other new aspects, as well as reorganize priorities to deal with the day-to-day operations while the foundation was in the process of hiring a new executive director.

We have been so greatful for the wonderful patience and kindness shown to us by you, our users and community, as we work through these growing pains. Keep submitting listens and let’s grow together!

– your ListenBrainz Team

3 thoughts on “Growing pains: An update on the ListenBrainz service status”

  1. Thanks for the detailed update and for being transparent about what’s happening behind the scenes. It’s reassuring to know that listens are still being stored safely. Hopefully these growing pains lead to a much more robust ListenBrainz

  2. Much <3 to the LB team!! We got this!

    Edit: Oh, this is quite a good opportunity to plug Alistral, a pretty sweet tool made by a community member. It will generate stats etc for your account, even if the LB servers haven't run them, as long as you keep submitting listens. It is a command line application (scary) but it has instructions for newbies: https://rustynova016.github.io/Alistral/

Leave a Reply

Your email address will not be published. Required fields are marked *