Rethinking Database Programming

14 hours ago (acadia.engineering)

The issue with defining schemas in a non-SQL programming language is they always lag behind what the underlying database can do. Sure, your ORM-like framework can define basics like primary keys and maybe uniqueness constraints, but can it define partitioning schemes, compression methods or more advanced constraints?

Look at all the features supported here:

https://www.postgresql.org/docs/current/sql-createtable.html

And then consider that other databases have even more. If you manage your schemas in code then you lose access to all of those, and will eventually need to write SQL anyway.

For queries it isn't such a problem, especially if you have a nice compiler. However, I recently lost faith in SQL wrappers/abstractions. The usual justification was that a lot of developers don't know SQL well, but LLMs are great at it. It's easier for the LLM to write SQL than some less familiar DSL. And SQL was written to be relatively easy to understand, especially if you do things like use CTEs and views correctly it should be possible to factor logic out to make even complex queries understandable.

The question for frameworks like Acadia is really: assuming I am fluent in SQL and know every feature of my database, what does the framework buy me? Because that's the perspective an LLM comes to it with.

  • There's a lot of benefit in these systems, though there's rough edges and I agree about the basics like PK's and uniqueness.

    I've been using Ormin [1] in Nim which works by parsing the SQL tables and uses it to compile time check queries:

        # Multiple joins with pagination
        let page = query:
          select Post(title)
          join Person(name) on author == id
          join Category(title) on category == id
          orderby desc(post.creation)
          limit 5 offset 10
    

    I think that's better since defining SQL should be the source-of-truth for the DB and the code. ORM's always ended up causing trouble in my experience.

    Things like indexes, defaults, partitions, etc generally aren't expressible in code without a lot of kludges. Then each DB engine have pretty different rules, syntax, etc for tables.

    However having the queries compile time checked, type conversions handled, and the nuances between SQL query syntax handled is rather nice. As you mention it's a much easier subset.

    1: https://github.com/Araq/ormin

    • Just learn SQL. I believe all these SQL replacement layers are just because people don't like SQL and don't learn it, so they learn a training wheels version of it that will cripple their ability to grow because it's simplifications remove expressiveness that caused SQL to be more complex to begin with.

      Just learn SQL, it's not that hard. A lot of very very smart people put a lot of effort into it. It's very good. The things that are annoy you about it are often there because of something you don't yet even realize is something you need to be aware of, or because your fundamental understanding of things is just wrong or incomplete.

      8 replies →

  • Agreed, that's why I chose to implement a simple ORM for my language's multi-platform database library. It has a mandatory 'id' column, for simple updating and deleting, but table creation and complex queries are done in plain SQL.

  • A core idea of the relational model is to seperate the logical model from the physical layer including optimizations, indexes etc.

    So it makes sense to only expose the logical model at the ORM layer.

    The problem comes if you want to define the database schema through the ORM layer, rather than just represet it.

    • Isn't SQL already a logical abstraction language over a "physical layer"? I'm not updating indexes or deciding when to flush or fiddling with MVCC when I write SQL

      2 replies →

  • > Look at all the features supported here:

    > https://www.postgresql.org/docs/current/sql-createtable.html

    Unironcally, yesterday i was vibe-coding a small app for personal use using Django and was quite shocked to discover that Django's orm does not support something as simple as specifying a database schema other than the default "public" one out of the box.

    You either have to add options specific from libpq:

        DATABASES = {
            "default": {
                "ENGINE": "django.db.backends.postgresql",
                "NAME": "mydatabase",
                "USER": "myuser",
                "PASSWORD": "mypassword",
                "HOST": "localhost",
                "PORT": "5432",
                "OPTIONS": {
                    "options": "-c search_path=myapp,public",
                },
            }
        }
    

    Or you have to do it from the postgresql side:

        ALTER ROLE myuser
        IN DATABASE mydatabase
        SET search_path = myapp, public;
    
    

    It's not ergonomic at all.

By now I stopped counting the attempts to replace SQL.

There is a lot of valid critic for SQL and I would be very happy if some things would have been designed different.

OTOH the architecture and mathematics behind relational databases are simple, composable and stood the test of time more than most other designs, methodologies or approaches to software development.

Though SQL can be improved, even with my average SQL skills I never had trouble getting information out of a database and fancy stuff like window functions make to my understanding even standard SQL Turing complete.

SQL has the native database support, for most companies the data and the database will outlive any specific application or even the whole ecosystem of a programming language/platform (Visual Basic, Visual FoxPro, Python 2, ...)

Further, we have fantastic books, knowledge, ORMs, query builders and a gigantic ecosystem in tools for SQL and SQL databases.

Acadia might be brilliant from a technological point of view, but it does not matter, because it does not look like a big enough improvement compared to SQL that it seems worth to invest in it. I will rather improve my knowledge of standard SQL or my knowledge for a specific relational database.

Finally Acadia does not really seem to raise the bar compared to other ORMs/Query builder. I get that from a FP point of view map/filter are nicer than a SELECT ... WHERE, but at some point in the projects I participated one would end up interacting directly with the database anyway, and at that moment I am back at SQL, so what did I gain?

  • SQL has one flaw: The verb should come last. So, "FROM users WHERE id = 1 DELETE" or "FROM users WHERE email = 'foo@example.com' SELECT id". That'd cut back on some accidental "oops I dropped the whole table" because I submitted a delete query before writing the where clause.

    Other than that, it's perfect, no notes.

    • The part that has stood the test of time and genuinely seems to carve reality at the seams is the query part. The data definition and data manipulation parts are just ok.

      2 replies →

I'm wary of languages that seek to own the database. In particular, the claim "Coexist with SQL" seems a bit suspect given that e.g. sum types have a custom binary encoding, which likely makes them difficult to interop with from other languages. This makes the claimed interop with other languages really more of a temporary stopping point towards full Acadia adoption rather than a viable long-term equilibrium, unless you e.g. eschew using sum types. (I also suspect that trying to natively support sum types can lead to a kind of FP-equivalent of ORMs' impedance mismatch. The ways I model data with relational logic can be pretty different than the ways I model data with algebraic datatypes and I wonder if trying to force fit the latter into the former doesn't lead to the same problems as force fitting objects into relational logic).

This makes the database closer to something that Acadia compiles to, rather than something Acadia sits on top of. From my own developer experience this feels off, because I generally expect the data layer to be king and application code to revolve around that, rather than having data representation created in code and the database created off that (this is why I also dislike things like ORMs).

In general I view databases as usually having more longevity than application code, especially as you accumulate more data over time. For serious production applications, the database often outlives multiple rewrites of the production application.

I suspect though my concerns are overall rather minor. The ergonomics of the language itself seem enjoyable. Acadia seems like it would be great as an embedded DSL. It's a bit unfortunate that it currently seems coupled to creating an HTTP server. I think that Acadia has greater ambitions beyond just the database, as evidenced by creating a binary web connection with frontend Elm code to presumably obviate the need for encode-decode layers. It seems like Acadia is meant to be a stepping stone towards a closer frontend-backend fusion. But I agree with mjaniczek that something like Lamdera seems a better fit for that.

But given how early Acadia is, I'm still very excited for where it goes. What I've listed is surmountable and I also feel that often a closer frontend-backend fusion might be worthwhile.

  • I think, the reality is SQL being simply to old to coexist with a web app use case. All the nice things that article talks about are not possible to nicely integrate with SQL. Current development is done by either writing SQL by hand or by letting ORMs to autogenerate it. Both feel bad because of how bad SQL is. But there is no other option. I hope https://substrait.io/ will gain traction and will be supported natively by databases

The main arguments presented are that databases do not support modern types.. And that this system replaces tested authentication systems by emailing a UUID in plain text?

It does mention UInt64 which is not a modern type and as far as I know is supported by every database.

It also compiles to SQL but it isnt clear where the advantage comes from other than using a different syntax to do things.

This looks reasonably interesting, and Evan is extremely thoughtful about design; I know he’s put a huge amount of work into this.

Personally, I’d be very cautious about adopting closed-source software with such a restrictive license as part of an application, especially given the context of Elm’s trajectory. When Elm went through breaking changes or regressions, or was not worked on publicly for years, users had access to the source and the right to modify it. With Acadia’s licensing, you’d be stranded.

  • On the other hand, with Elm there was no correlation between adoption and funding for development. With Acadia, he's trying a different funding model, so that might mean better support for both Acadia and Elm.

    • The Elm project forked into a bunch of different Elms because Evan basically abandoned / killed it. Then he got more interested with this project. What’s to say that won’t happen again?

      8 replies →

I don't see anything special here. Haskell has had stuff like this for more than a decade, Selda is probably the one closest to Acadia: https://valderman.github.io/selda/

Despite their claims, this is not substantially different from ORM platforms in many languages.

I'm actually most excited about the new funding model: https://acadia.engineering/license/faq

I won't be able to use Acadia at work, and I don't have the risk tolerance to use it for personal projects, but I'm looking forward to seeing how/if this model pays the bills. Can it compete with more liberally licensed code?

It seems like this is a few things:

1. An Elm-like programming language that lives in .db files

2. A compiler from this language to strongly-typed database procedures in a target backend language

This has more in common with a semantic layer than an ORM.

What you gain is a shared language that connects the table definitions (say a SQL migrations folder) and your API language (often handwritten SQL). This can be type checked and optimized for you.

But for me the big question is what functionality do you lose? Can I express everything that PostgreSQL can?

So this is capable of turning a one-liner of SQL into six lines of barely readable code?

  • It seems that is the price you pay for the power to turn a 600-line nightmare SQL query into 60 lines of barely readable code.

    • I would like to see that example then.

      I’m all for improving on SQL, but this syntax does not even solve the dangling comma issue as far as I can tell from the example.

    • I'd rather take the 600 lines of SQL, provided it's not dynamically constructed. SQL is a very high level language, it's fine IMO.

  • SQL is a horrible language. I’d gladly program in something composable like Elm.

    • As a programming language? Sure. As a way to work with relational data? It may be my favorite "language" across all domains because of the terse beauty. I am a self-taught, no CS coder but SQL is the one place where I feel like I get all the math I should know.

      An opinionated, possibly hot take would be to call SQL "A more elegant weapon of a civilized age".

      2 replies →

    • Maybe so, but my father in law, who is a salesman and knows nothing about computers and programming still knows SQL.

      SQL is a horrible language in the same way Excel is -- programmers hate it but the what makes it a horrible programming language to developers is what makes it accessible to non programmers.

Some interesting features here:

- sum types/ADTs have been long missing from database data modeling and this is welcome change. It's not entirely clear to me how the migration strategy here will work with things like removing a variant, etc.

- first class enforced RLS - this seems like a fantastic way to ensure safety/security guarantees. Secure by construction is always preferable to bolt-on security controls.

- composability with a strong module system. I think this will work well in ensuring large schemas can evolve over time. I wonder if there will be package manager in the future.

Having reusable functions and pipelines compiling to SQL sounds amazing. (EDIT: and sum types!) Will want to try this out on some side project later.

Although for my Elm + backend needs I feel like I still prefer Lamdera: https://dashboard.lamdera.app/ - WebSocket communication and being able to push new data to clients immediately instead of juggling HTTP endpoints and the client having to pull/refresh. `sendToBackend`, `sendToFrontend`, `broadcast` are a great primitive.

I have a very long history with language interfaces to databases.

- As a grad student in the 80s, I read a lot about "database programming languages", which aimed to provide persistence and query capabilities to conventional programming languages, in a seamless way.

- The next step to putting those ideas into practice: Participated in a research project on adding database capabilities to a programming language (anyone remember Ada?)

- I designed and developed most of the modeling and query language features of one of the major object-oriented database systems, back in the early 90s.

- I also designed and contributed to a SQL interface to our OODB, as well as an ORM, taking our model and query language, and mapping it to SQL.

- Turned down an offer from a software giant of the late 90s, to add database capabilities to one of their main languages, (basically bringing to their language what I had built at the OODB company).

- Designed and built a Java ORM (late 90s).

And after working on this stuff for something like 20 years, I concluded that it's all misguided. For all of its ugliness and weirdness, SQL was designed to address a certain set of requirements, and has succeeded wildly. New database programming languages face huge problems of acceptance, and needing to solve the exact same problems that SQL handles now. (This was easier 30 years ago since it was still early days for SQL. Now it's basically impossible.) ORMs are a terrible idea, in the "now you have two problems" category. Not only do you need to write high-performance queries, but you have to get your ORM to actually issue those queries. (Yes, ORMs have escapes to raw SQL. The existence of these escapes proves my point.) And schemas change, and the mapping to your language model has to change, and it's a mess.

Just use SQL. It's the right tool for the job it was designed for. Use a database driver to integrate with your language. It's just not that hard.

  • The thing I’ve never understood is why SQL itself is not the target of attack. There’s already an inherent language abstraction with the planner; Postgres in theory could be the JVM with any number of languages implemented on top. Including a language that lends itself to composition and auto generation of PL functions.

    ORMs are fundamentally difficult because of the mapping problem, but SQL code builders should be trivial. Auto-generating and exposing every DB functionality as a type-safe $LANG function should be trivial. Instead, they’re also accidentally difficult because building SQL is difficult.

    Outside of SQL, you’ve got datalog… and that’s about it. And I guess whatever horrors the NoSQL crowd keeps coming up with

  • Given your experience, what is your opinion on stored procedures?

    I love them, think that what can be done in the database should stay in the database, and many of these abstraction on top are all ways to avoid just having to implement them.

    And the main reason, DB portability, seldom happens in reality, most product die still using the database they were original created with.

Looks very nice. Last year I took up rust, coming from c++, and some of the modern features rust brings are just so nice to have (even something as simple as not having to forward declare a class).

This year I started working with postgres and you just can't help but notice how sql is coming from the c-Era of programming. Having better and more modern ways to express my queries would be great to improve correctness and performance.

  • > …can't help but notice how sql is coming from the c-Era of programming. Having … more modern ways to express my queries would be great to improve correctness …

    SQL is based in pure mathematics: set theory, relational algebra.

    The process of applying mathematical rigor to your database design to prove correctness is referred to as normalization.

    I don’t mind criticisms like “It’s old, yuck”, but criticisms like “it’s not correct” mean you haven’t studied or applied the mathematical underpinnings of sql.

    • Syntax aside, programmers and mathematicians have a very different view on how things should be done.

      Programmers look at data and see opportunities for running a pipeline of transformations (map/filter/...). And they tend to write their SQL like this as well. Or use something like Linq or one of the various pipe syntax SQL extensions.

      I would say that this is a major reason why there is this sentiment of "SQL is yucky" by developers. The mental models just don't match.

      2 replies →

    • This isn’t talking about correctness of SQL. It’s talking about correctness of queries.

  • It is older than C. It is based on COBOL era idea of structured English as a computer language. There are better alternatives, e.g. Datalog.

    • I'm curious if you've personally used datalog in any projects. I've written some prolog, but haven't ever worked with datalog.

      Minigraph looks promising for some introductory goofing around.

    • These kind of comments don't age well in the days of AI programming using English.

The idea of treating database programming more like regular programming is nice. I'm just not sure how much complexity this actually removes versus moving that complexity somewhere else

I was hoping for an alternative to PLSQL or stored procedures. But this isn’t about „Database Programming“, it’s a SQL replacement…

  • It isn't that bad, at least for those of us that like Ada, and feel at home on SQL Developer.

Reading this, I mistook it for a slightly different idea: using these functional languages directly inside the database process, avoiding SQL altogether.

I've wanted to try that out with e.g. Roc and a reimplementation of SQLite's on-disk format. (Of course, that's a non-starter for production use, but it could be an interesting experiment to see what that programming model was like.) The database would become kind of like a library you use to build your tables and queries with.

Also, thank you for calling it a 1+n query, not an n+1 query ;)

I agree with the premises, but the result proposed here doesn't look like anything I would like to use unfortunately. Even just looking at a glance you cannot see what it's doing and what each part means.

Oh man. If this lobste.rs comment is correct about the subscription terms then this feels like a really hard pill to swallow: https://lobste.rs/s/ykq7ym/rethinking_database_programming#c...

Still might be viable, but would be tricky to sell.

> SUBSCRIPTION TERMS

> This license is subscription-based and will remain valid only for the duration of your active subscription. Upon expiration or termination of your subscription:

> a) Your rights to use the Software will cease; b) You must uninstall and stop using the Software; and c) You may lose access to any data or content created with or stored in the Software.

  • On the other hand, norms in software right now are that suckers build and maintain software for free + "the love of the game should be enough for anyone", so it's shocking when people break the norm.

Needs proper docs

stuff like "The endpoint keyword" just gets a mention on the front page/readme with no further detail

Is this at all similar to LINQ in C#? I never used it, but I'm vaguely aware of it being a functional approach to querying an RDBMS.

  • From what I seen (not an expert). It’s mostly sql with a c# flavor and auto translation to native type.

Hmm. I've skimmed the article. It looks to be another ORM/FRM type thing. There are many issues with such things, but for me the most troubling is this: in most systems (obviously...it depends) you don't want to wind the database around the axle of any one software component or language. Having the data separate from the code, and defined/managed with a language that suits data management is a feature not something to be designed out. My hunch is that people who come up with these "solutions" fail to realize this. They then condemn everyone using their layer to endless hair pulling trying to figure out "what SQL did it make from that?" and "how do I make it do this SQL?".