ArangoDB, or when relationships matter more than tables

Created: May 18, 2021 Updated: Sept. 29, 2026

In one project I used a graph database in production for the first time, a few words on why and whether it was worth it.

The project is about digital cultural heritage, we collect digital copies of objects from various institutions, libraries, museums, archives, and each has its own data format, so the integrations alone take a lot of time. The main database is Postgres, search goes through Elasticsearch, and that works well as long as we're dealing with single objects.

But objects don't live on their own. A painting has an author, the author had a teacher, the painting was shown at an exhibition, the exhibition was held in some city, and in that city lived another person we have a letter from in the archive. We wanted people to be able to follow these links, and in tables that meant a pile of join tables and recursive queries nobody wanted to read afterwards.

ArangoDB keeps documents and the graph in one place, an object is simply a JSON document and the links are edges between documents. A query like “show me everything at most two steps away from this person” looks more or less like this:

aql
FOR v, e, p IN 1..2 ANY 'persons/matejko' GRAPH 'heritage'
  FILTER v.type IN ['artwork', 'exhibition']
  RETURN DISTINCT { title: v.title, via: p.edges[*].label }

Instead of a dozen lines of SQL with CTEs you get a few lines you can actually read.

The costs

A second database is a second thing to maintain. Postgres is the source of truth and we build the graph from it with Celery tasks, you have to make sure the two don't drift apart, back up both, and new people on the team have to learn AQL.

I'd choose it again, because here relationships are the heart of the product. If the graph were just an add-on to an ordinary application, recursive queries in Postgres would do just fine.

#arangodb #grafy #postgresql

Machine-translated from Polish (original).