← Back to context

Comment by refset

6 years ago

I don't have experience with either RecallGraph or ArangoDB but I have been following the project for a few months now - kudos to the team for releasing 1.0!

Naturally there are a few other open source technologies exploring the "temporal graph" space, including TerminusDB [0], which uses a git-like branching history model for data collaboration, and Crux [1], which supports bitemporal history for transactional workloads (full disclosure: I work on Crux directly). Datahike [2] is also worth checking out if you've not seen it. Both Crux and Datahike are heavily inspired by Datomic.

However, if the history-aspects of your data model end up being more complicated than the graph-aspects of your data model, then you may well be better off using SQL temporal tables as implemented by Teradata / SAP HANA / DB2 etc.

[0] https://terminusdb.com/

[1] https://opencrux.com/

[2] https://github.com/replikativ/datahike

Like you, I don't have any experience with RecallGraph, but I have been following Arango closely (I work with TerminusDB). First impression is that it looks very interesting and going in a similar 'like Git but for data' direction as Terminus. I'd be somewhat concerned about building on Arango - without your own storage layer, some of the engineering challenges will be very tricky.

When we first built Terminus, we used Postgres as a store, but found that it was too slow for the types of queries we wanted to run. After a HDT detour, we built our own in the end(https://terminusdb.com/blog/2020/04/14/terminusdb-a-technica...)

I think graph dbs (crux, terminus or recall) are very natural places for revision control.

Of these Crux, looks quite interesting. How well can it scale? Can it handle billions of documents?

For us, history is mostly about being able to audit a record and understand it history, and occasionally undo a mistake. Don't need or want to go with things like SAP.

  • Crux is designed to scale directly based on how RocksDB (or LMDB) performs on a single node, in terms of: sustained ingestion throughput, KV seeks/sec, and the sheer quantity of KV data that can be supported on an array of local SSDs (i.e. easily many billions of small docs). At a higher level this means that point-in-time queries will maintain good performance regardless of how much history is stored, thanks to Z-order indexing [0], and the query algorithm only requires very modest amounts of memory because the KV indexes are lazily streamed out of Rocks and processed tuple-by-tuple (though having more memory is always going to speed things up!).

    Beyond the scope of a single Crux node, horizontal read scaling comes for free due to the transaction time model of history (i.e. you can spin up N identical nodes to service all manner of wholly unrelated use-cases with consistent reads).

    [0] https://en.wikipedia.org/wiki/Z-order_curve