Acoustic fingerprints: Is closed source OK?

Ever since my post about TRM hitting its limits, I’ve been in discussions with a reputable company who has offered to let the MusicBrainz community use their fingerprint server in exchange to a free license to the MusicBrainz live-data feed. While I am not ready to reveal who this company is, I do feel that … Continue reading “Acoustic fingerprints: Is closed source OK?”

Ever since my post about TRM hitting its limits, I’ve been in discussions with a reputable company who has offered to let the MusicBrainz community use their fingerprint server in exchange to a free license to the MusicBrainz live-data feed. While I am not ready to reveal who this company is, I do feel that I can trust these folks — this is not the first time we’ve chatted.

The straw-man deal that we’ve put together makes sense for MusicBrainz and this company. Unlike MusicBrainz’ relationship with Relatable, this relationship would be more balanced. Plus, we would not have to maintain the server ourselves. All around I feel good about this proposed deal. There is just one little snag.

They are uncomfortable with open sourcing their client.

While I am not an open source license Nazi, I have received tons of complaints about MusicBrainz using a technology that is not fully open. As a matter of fact, my most unpleasant dealings with the general public have been on this point (that the TRM server is closed source). And I’ve had unreasonable people shout unreasonable things at me over this point. Quite frankly I am not really interested in having to defend my position on this any further, but I fear that not having a working fingerprint solution may be more of a hassle than having to defend a closed source solution.

So, my question to you is this:

  1. Do you value having access to a fingerprint solution as part of MusicBrainz more than MusicBrainz being an end-to-end open solution?
  2. What arguments can we make for having this company open source their client, as Relatable did? I’ve argued the standard open source arguments and I think that there is still a small chance that we can persuade this company to open up. I need to construct a better argument and perhaps meet with them in person to hash this out further. What things should I argue?

Please keep your idealistic everything needs to be open arguments to yourself. I simply won’t bother reading them or responding to them. I really care to see if we can find a balance where we can maximize the value that MusicBrainz presents, even if it means compromising our values slightly. If you’re not ready to make a balanced argument, then please don’t.

NOTE: If we were to start using a closed source fingerprint solution, nothing else would change. None of the existing licenses for MusicBrainz would change. So, keep your pants on and stop frothing at the mouth.

MetaBrainz milestone: My first paycheck!

About a month ago, the MetaBrainz Foundation signed up its second customer (let’s call that customer Mystery Customer #2). In light of this and our current finances our Board of Directors approved for me to make this payment my very first paycheck. Early today the first payment from this customer hit our PayPal account and … Continue reading “MetaBrainz milestone: My first paycheck!”

About a month ago, the MetaBrainz Foundation signed up its second customer (let’s call that customer Mystery Customer #2). In light of this and our current finances our Board of Directors approved for me to make this payment my very first paycheck.

Early today the first payment from this customer hit our PayPal account and I cut myself the very first MetaBrainz paycheck ever! While this check for $2400 is measly pay for a years worth of work, its a start. With more customers on the horizon I can hope for a reasonable paycheck in the next year. And me getting paid is good since I can stop hunting for contract work to pay the bills — this will allow me to focus more of my time on MusicBrainz. And everyone knows that we need to implement more features, right?

A new Picard, a new libmusicbrainz, a new libtunepimp and now a paycheck! What a week, and its only tuesday!

P.S. I can’t reveal who our customers are, since they are building new services based on our data. These services have not been announced yet and us talking about who these customers are, would tip their hands. As soon as these companies go public with their services, I’ll be sure to tell everyone who they are.

Technorati Tags: , ,

Picard 0.5.0 released!

After too many months of tinkering the latest stable release of Picard has been released: picard-0.5.0.tar.gz (Linux tarball) picard-setup-0.5.0.exe (Windows installer) Big thanks goes out to Lukas Lalinsky for fixing many bugs and creating the Windows installer. Also, many thanks to everyone who helped work on the Norwegian, German, French, Russian and Slovak translations! Technorati … Continue reading “Picard 0.5.0 released!”

After too many months of tinkering the latest stable release of Picard has been released:

Big thanks goes out to Lukas Lalinsky for fixing many bugs and creating the Windows installer. Also, many thanks to everyone who helped work on the Norwegian, German, French, Russian and Slovak translations!

Technorati Tags: , ,

Continue reading “Picard 0.5.0 released!”

libtunepimp 0.4.0 released

libtunepimp 0.4.0 has also been released! Long overdue since there hasn’t been a libtunepimp release in over a year. This version of libtunepimp supports a complete plug in architecture for different media formats, is fully UTF-8 compliant, supports ID3v2.3 & ID3v2.4 tags and fixes a number of old bugs. Version 0.4.0 is not compatible with libtunepimp 0.3.0 — so, if you have an application that uses libtunepimp 0.3.0 you may want to wait to install the 0.4.0 version. This release contains the following files:

Technorati Tags: , ,

Continue reading “libtunepimp 0.4.0 released”

libmusicbrainz and libtunepimp test releases

On the excessively long road to the next stable Picard release I need to do a stable release of libmusicbrainz and libtunepimp. Both of those are now complete with all the bug fixes that I want to get into these before the next release. If you run Linux or Mac OS X, please take a moment to download, compile, test and install these two tarballs:

  1. libmusicbrainz-2.1.2-pre1.tar.gz (Changelog)
  2. libtunepimp-0.4.0-pre6.tar.gz (Changelog)

Please note that the independent python bindings for libmusicbrainz have been wrapped back into libmusicbrainz itself and can be found in the python directory of the tarball. libmusicbrainz has a new example program called getrels that shows how to retrieve AR links via the web service. libtunepimp has undergone major changes and now sports a plug in system that will make it easier to add more formats.

Unless these tarballs have major issues, I plan to update the version numbers to 2.1.2 and 0.4.0, respectively, and release them in the coming days. Then I’ll focus on getting stable release of Picard out the door.

Technorati Tags: , , ,

Server mini update

We just updated the server with a set of minor bug fixes. Mostly formatting and appearance bugs have been fixed and new preferences to control how the top menu works have been added. That’s it — happy brainzing!

We just updated the server with a set of minor bug fixes. Mostly formatting and appearance bugs have been fixed and new preferences to control how the top menu works have been added. That’s it — happy brainzing!

Server updated finally!

The server has been updated – this release includes a redesigned navigation system, changes to the Autofix (a.k.a. Guess Case) tool and Album Editor. Track durations can now be edited and a number of smaller bugs have been fixed. A big thanks goes out to g0llum, lukz, matt and djce for making this release a … Continue reading “Server updated finally!”

The server has been updated – this release includes a redesigned navigation system, changes to the Autofix (a.k.a. Guess Case) tool and Album Editor. Track durations can now be edited and a number of smaller bugs have been fixed.

A big thanks goes out to g0llum, lukz, matt and djce for making this release a reality. Now that we have this pesky re-org release out of the way, we hope to bring you several more releases before the end of the year.

Continue reading “Server updated finally!”

Acoustic fingerprinting at MusicBrainz: The Future

My last post on current state of the TRM fingerprinting solution got quite a bit of response — I was quite amazed by it really. Personally, I think people still put too much emphasis on TRM and what role it plays within MusicBrainz, but without me providing a new tagging solution there aren’t any concrete … Continue reading “Acoustic fingerprinting at MusicBrainz: The Future”

My last post on current state of the TRM fingerprinting solution got quite a bit of response — I was quite amazed by it really. Personally, I think people still put too much emphasis on TRM and what role it plays within MusicBrainz, but without me providing a new tagging solution there aren’t any concrete points to discuss.

Given the feedback I’ve gotten, I’d like to state a reformulated vision with regards to acoustic fingerprinting and tagging here at MusicBrainz. The two points that have received the most feedback concern acoustic fingerprinting and downloading large index files in order to use the tagger.

Acoustic fingerprinting: Since so many people professed their love for TRM and acoustic fingerprinting in general, we will do the following things:

  1. Keep TRM alive.
  2. Work to create an open replacement for TRM. See the musicbrainz-devel mailing list for discussion on this topic and if you would like to help out. The founder of Tuneprint has recently volunteered to help build this new solution and I expect that his presence in this project should stir things up a bit.
  3. When #2 is operational, we will start a gradual migration to the new server. TRM is not going away tomorrow! Got it?

The obvious problem is if #2 does not come to fruition — if you care about TRM and acoustic fingerprinting here at MusicBrainz, you should go check out the discussion on the devel mailing list and lend your hand. If it doesn’t come about and the TRM server stops being useful, then we’ll eventually turn the TRM server off.

Picard & large indexes: The Picard tagger with Lucene support will progress as planned — the only change so far will be that I will provide one machine for use as a centralized lookup server that will not require you to download the massive text index. However, I expect that Picard with Lucene will be a popular tagging tool, and that the server will get overloaded and slow in the space of a few months. Given that, we’ll have complete indexes available for people to download.

I predict loads of people will opt to download the text index since a 250Mb download will be a lot faster than trying to tag their 10,000 file collection on an overloaded server that performs 10 lookups per minute for them.

Thanks for all the feedback!

UPDATE: PLEASE stop telling me how much the large index would cramp your style and how much the fingerprinting has saved you. I know!

General update: What's up with TRM??

This general update is way overdue — a lot of things have been happening behind the scenes and its time to let everyone know where things in the MusicBrainz world are headed. I’ll start off with TRM, since that is hot discussion topic on the musicbrainz-users mailing list right now. The TRM (TRM’s are acoustic … Continue reading “General update: What's up with TRM??”

This general update is way overdue — a lot of things have been happening behind the scenes and its time to let everyone know where things in the MusicBrainz world are headed. I’ll start off with TRM, since that is hot discussion topic on the musicbrainz-users mailing list right now.

The TRM (TRM’s are acoustic fingerprints that MusicBrainz uses to identify music tracks) server is constantly overloaded and can only handle a database size of about 2.2Gb before it crashes. To prevent crashes, we prune the database where we throw out the least used TRMs, which implicitly discards work that our users have done. Not good. In order to make the TRM server perform at some reasonable level of performance, the entire database needs to be kept in RAM. Thus our server has 5GB of RAM and it still can’t keep up. The fact that this problem hasn’t reared its ugly head to the public, is a testament to Dave Evans’ skill in keeping the TRM server ticking.

Furthermore, TRMs have shown themselves not to be as unique as we would’ve liked. For example, take a look at the TRM’s with at least 5 tracks report: 4400 pages (!) of TRMs that I would consider to be sub-optimal. One example TRM (non silence on page 2) has 104 tracks associated with one single TRM. Given this, TRM is not some sort of magical solution that with great authority tells the tagger what metadata to apply to a track. Instead, its best to think of TRM as a system that lets you guess which few dozen tracks a file could be matched to — there is a lot of logic in the tagger that makes up for the shortcomings of TRM.

Thus, TRM has two major problems: its not accurate enough and it doesn’t scale well to the size that MusicBrainz has grown to. The system still functions but I expect it to start breaking down and becoming of less use over time. We have the following options:

  1. Find a replacement for TRM: Relatable doesn’t seem to be in business anymore, or at least they are in deep hibernation. No other companies that I have approached were interested in sharing their technology with MusicBrainz. (For the record, I’ve tried with 3 companies, including a couple of on-site visits in Europe).
  2. Create our own TRM solution: This is an very large endeavour — at least a year if not two, of hard work. I’d rather work to improve MusicBrainz itself, rather than hacking on acoustic fingerprint software.
  3. Throw more resources at TRM: We’re still lacking the funds for more resources, and the same argument in #2 still applies.
  4. Do something else: Find some technology that can replace TRM.

Given my babbling about Lucene, I think its a foregone conclusion that #4 is the way to go. Sometime this fall, I will release a Picard tagger with a lucene text indexing engine to replace the current MusicBrainz Tagger. The benefits of this new tagger will be:

  1. It will distribute the load on the server, since currently a large chunk of the server load goes to supporting tagger users. And a large chunk of tagger users never really contribute data to MusicBrainz or make cash donations to support the project. So, moving that traffic off the main server will allow people who want to edit/vote on the data focus on their work.

    Given that most files in the wild nowadays have some metadata, a text index will work well. Lucene is great at taking crappy data input and coming up with something useful. If TRM gets us into the ballpark and then additional heuristics do the final leg work, Lucene will give us a much better guess to start with than TRM ever did. Thus, overall tagging quality will improve greatly.
  2. A lucene tagger will work much faster than the TRM based tagger ever was. 2-5 seconds per track was not unusual given TRM — with Lucene we’ll see 2-5 tracks per second, if not much faster.
  3. Since we will no longer have to decode files to identify them, it will be easier for us to support new formats. Its less work overall.

This approach also has the following downsides:

  1. It will no longer support identifying completely anonymous files. Files that have no id3 tags and are named test1.mp3, test2.mp3 will simply not stand a chance at identification. I realize that there is great romance associated with this concept, but in reality most people have files that have some metadata in them, and thus will stand a good chance of being identified.
  2. You will need to download a 250MB Lucene index to tag your collection. This is a pretty big hurdle, but if BitTorrent can routinely help people download 650Mb movies off the net, it should help us download distribute our search indexes. After the first release of a Lucene enabled Picard, we will investigate P2P searching methods that will allow people who have no index to use some other people’s indexes (if they allow that).

So, the roadmap for this looks like this:

  1. Release picard 0.5.0 in the next few weeks and start putting it on the main page as an alternative to the MB tagger.
  2. Release picard 0.6.0 with full Lucene support and offer that as the main tagging solution for MB.
  3. When the TRM usage drops because of adoption of Picard 0.6.0, we will start phasing out TRM.

There you have it — thats the current happenings on TRM and how we hope to solve the problems that it presents us with.