Comment by Bedon292
6 years ago
I had recently been thinking about something like this. A project I have been working on has lots of relational data in postgres, and keeps the history of all the data and relationships, but its so relational a graph is something we were considering for the next version. But the history is important. So this might actually fit the bill.
Has anyone had any experience with this yet?
I don't have experience with either RecallGraph or ArangoDB but I have been following the project for a few months now - kudos to the team for releasing 1.0!
Naturally there are a few other open source technologies exploring the "temporal graph" space, including TerminusDB [0], which uses a git-like branching history model for data collaboration, and Crux [1], which supports bitemporal history for transactional workloads (full disclosure: I work on Crux directly). Datahike [2] is also worth checking out if you've not seen it. Both Crux and Datahike are heavily inspired by Datomic.
However, if the history-aspects of your data model end up being more complicated than the graph-aspects of your data model, then you may well be better off using SQL temporal tables as implemented by Teradata / SAP HANA / DB2 etc.
[0] https://terminusdb.com/
[1] https://opencrux.com/
[2] https://github.com/replikativ/datahike
Like you, I don't have any experience with RecallGraph, but I have been following Arango closely (I work with TerminusDB). First impression is that it looks very interesting and going in a similar 'like Git but for data' direction as Terminus. I'd be somewhat concerned about building on Arango - without your own storage layer, some of the engineering challenges will be very tricky.
When we first built Terminus, we used Postgres as a store, but found that it was too slow for the types of queries we wanted to run. After a HDT detour, we built our own in the end(https://terminusdb.com/blog/2020/04/14/terminusdb-a-technica...)
I think graph dbs (crux, terminus or recall) are very natural places for revision control.
Of these Crux, looks quite interesting. How well can it scale? Can it handle billions of documents?
For us, history is mostly about being able to audit a record and understand it history, and occasionally undo a mistake. Don't need or want to go with things like SAP.
Crux is designed to scale directly based on how RocksDB (or LMDB) performs on a single node, in terms of: sustained ingestion throughput, KV seeks/sec, and the sheer quantity of KV data that can be supported on an array of local SSDs (i.e. easily many billions of small docs). At a higher level this means that point-in-time queries will maintain good performance regardless of how much history is stored, thanks to Z-order indexing [0], and the query algorithm only requires very modest amounts of memory because the KV indexes are lazily streamed out of Rocks and processed tuple-by-tuple (though having more memory is always going to speed things up!).
Beyond the scope of a single Crux node, horizontal read scaling comes for free due to the transaction time model of history (i.e. you can spin up N identical nodes to service all manner of wholly unrelated use-cases with consistent reads).
[0] https://en.wikipedia.org/wiki/Z-order_curve
There are those who've started integrating this product into their stack. You can find them at the project's Gitter community forum at https://gitter.im/RecallGraph/community
Note that graph-like data can be represented quite easily in relational DB's including postgres, and general graph queries are made possible via the recursive-CTE feature in newer SQL standards.
If the graph is not the important part but want a serverless time-history stateful db check fauna.com
Maybe check out Cayley, Google's graph database which can use postgres as a backend?