Current ListenBrainz server issues/high load, Part 2

This follow-up post covers how and why our servers are still being impacted, an explanation and examples of what we consider “botting” in ListenBrainz – and an apology, because our last post had some baaad comms!

Firstly, we are still experiencing massive server load from an increase in listen submissions related to the BTS fan exodus from last.fm (not all bad! more on this below) – and we are now being hit by another influx of AI scrapers.

Our apologies to our users, new and old, as we continue to work to mitigate the extra load. Remain assured that listens are not being lost – our processing queue is lagging behind, so listens may take (up to) hours to appear.

LB radio and the popularity endpoints have been temporarily disabled, to reduce server load.

ListenBrainz server status/load

Our ingestion queue for the past 2 weeks.

The above screenshot shows our ingestion queue (listens being submitted directly to ListenBrainz) for the past 2 weeks, which usually sits around 0 with peaks at 10K. The highest peak during the last two weeks – the big mountain on the graph – is 484K. That is to show both that the listens are not lost and why they were lagging behind. “Your listen is now number 484,001 in line…”

Our import queue for the past 2 weeks.

The above screenshot shows our import queues (listens being submitted via our import function), in the same 2-week window. The values are “number of listens added to the ingestion queue per minute “. The purple line is lastfm imports. We usually sit around 1-2K, with peaks around 20K sustained for ~3 hours at worst. The peak in this graph is 180K – that’s 180K listens added to the queue that minute. These values have not been reducing at the time of writing, but fluctuating depending on time of day.

You can see these graphs, and the current length of the ‘queue’, in real-time on this Grafana dashboard (free account required).

Where this increased load is coming from

Our server issues are currently three-fold. One is good, the other two not so much.

  • Last.fm imports: This is the good problem! When a user syncs their account from last.fm, they’re sometimes importing 20 years worth of listens at a time. Every listen that is added to ListenBrainz passes through our listen queue, where our system checks the attached metadata, and tries to match it to the right artist, album and song. As you can imagine, if hundreds or thousands of users add their entire last.fm history at once, it’s quite the processing queue. This load is a sign that we are growing, and we don’t mind at all. Though we also don’t enjoy having to wait for listens to appear – if anyone wants to gift us some extra servers….
  • “Botting”/listen submission spamming: This is not good. A very small minority of BTS fan accounts are doing things like submitting a listen every few seconds, and/or submitting listens across multiple accounts at once, which has resulted in hundreds of thousands of extra listens being added to the queue. This goes against our code of conduct. Please scroll down to the end of this blog post for clarification on what behaviour that we consider botting.
  • AI scrapers: There’s not much to say here. The MetaBrainz foundation offers all of our data as freely downloadable datasets, but these AI companies don’t care – they set up scrapers to blindly hoover up the internet without respecting limits or ethical considerations. They go to great lengths to hide behind thousands of residential IPs, making them extremely difficult to detect and block (as blocking these IPs also blocks legitimate users). You can read more in our previous blog on the subject: We can’t have nice things… because of AI scrapers

Bad comms!

We – the MetaBrainz team – are very sorry for how our last post was written, and how it negatively affected so many people.

We are used to speaking to a small audience of our existing users, and we have a wonderful BTS/K-pop community in *Brainz already, which made us be too relaxed with our tone.

It is true that a small minority of fans recently started doing things like submitting a listen every few seconds, across many accounts, which has impacted our service for all users. But we <3 our K-pop community! So we never imagined that it would be thought that we are blaming all the fans of a band or even a whole genre. But that is not obvious from outside our small community and usual readership, and that is our mistake and something we should have considered before posting.

We are a small team, and most of us are very… technically minded. We do not have a comms team, and will probably never have one, so we ask for you to be patient with us.

On the other side of the coin, we had to filter a lot of abusive comments from the last blog post (both from ARMY and directed towards ARMY) and that is not the sort of community we are fostering here. We also are, and will continue to be, direct and up-front with our messaging, in accordance with our open ethos. If a group of users – even if it’s a non-representative subset of a larger group – are not respecting our rules and are causing issues for everyone, we will say so, and we will share how and why. We will endeavour to do this with tact and awareness (which we failed to do in our last post).

If that works for you, and the respect is shared both ways – welcome, new users! And if you are a BTS fan, congratulations on the new album release, and we look forward to seeing it continue to explode the global charts.

What we consider ‘botting’/bad behaviour in ListenBrainz

First, please note that when we say “botting”, we don’t mean that the listener is not a real human! It refers to using digital tools/code/devices to submit song listens that don’t reflect real listening.

We also want to acknowledge that this might be counter to how some people want to submit and store listens, and may not match everyone’s definition of ‘real listening’. We are sorry that our systems cannot cater to every approach. We think it’s awesome that the BTS community has their own app, B-CD, that caters to users who want to listen-max their love for BTS! We love to see it!

The following are (just some) examples of what we consider unacceptable in ListenBrainz.

  • Submitting listens you haven’t listened to: If you are submitting a 1 minute song as ‘listened’ every 5 seconds, or any other timeframe that is far shorter than the song itself, we are likely to disable your account. We have paused 100 accounts doing this, so far, with the ‘busiest’ submitting 6,000+ listens a day for a few days – a full song ‘listened to’ every ~15 seconds, day and night. This adds up.
  • Submitting the same listens across multiple accounts/devices: Our systems expect 1 song submitted listen to equal 1 song going into your earballs. Having multiple accounts/devices each submitting the same listen doesn’t get around this! If you are doing this we are likely to disable your accounts.
  • Doing both of the above at a large scale: “Ahhhhhhhhhhhhhhhhhhhhhhhh” – our servers.

tl;dr if your 1 listen is becoming 2 or more listen submissions in ListenBrainz, in any way, it’s likely that your account will be disabled. If you email us we are always happy to look at the situation and to reinstate accounts for people who agree to follow the ListenBrainz rules. Our apologies that we cannot proactively email everyone who we disable at this time, we are dealing with too many accounts to do so.

If you read all the way down here, thank you and have a lovely week – your MetaBrainz and ListenBrainz team.

4 thoughts on “Current ListenBrainz server issues/high load, Part 2”

  1. Unfortunately multiple submissions for the same listen do seem to be more common these days than I’ve previously noticed.

    On Last.fm I’ve found an artist with 224 listeners and 774k scrobbles, which certainly has the scrobble count inflated.

    It is however possible to have 2 legitimate listens within seconds of each other depending on the platform as listens on Spotify (via last.fm) record the scrobble timestamp as the start of playback and others (including Plex) record the timestamp as the end of playback of a track, and I suspect if you have both last.fm and Listenbrainz connected to Spotify, you’ll get 2 separate timestamps for the same listen. I purposely only scrobble directly to last.fm to prevent this from happening.

  2. This is somewhat off-topic but regarding the AI scrapers I do not understand why you allow AI code in *Brainz projects including Picard which is excessively using Github Copilot which uses stolen FLOSS code without attribution. You are supporting the very system that is mindlessly scraping *Brainz projects for endless profits. I kindly ask thinking over your current AI Policy and revising it to ban the use of unethical LLM coding tools in the projects. Thank you

  3. Thanks for the transparent update. I appreciate the team’s hard work in keeping ListenBrainz running despite the unexpected surge in traffic. Looking forward to the improvements!

  4. Brainz projects for endless profits. I kindly ask thinking over your current AI Policy and revising it to ban the use of unethical LLM coding tools in the projects. Thank you

Leave a Reply

Your email address will not be published. Required fields are marked *