← Back to context

Comment by uncle_kostya

8 hours ago

I'm curious about motivation - bounds checks for corrupted inputs seems like it would be one, but it also seems that fixing corrupted input handling in a C/C++ code base would not be too hard, and probably less of an effort? So why did you choose the rewrite?

And second, did you use any AI tools for the rewrite?

> fixing corrupted input handling in a C/C++ code base would not be too hard

The best programmers on the planet have tried and failed with this task for 50 years now, so I don't think this is true.

The main disadvantage of Rust right now is not supporting some more obscure platforms, but because mold wouldn't support them anyway I don't see that as a problem.

  • > The main disadvantage of Rust right now is not supporting some more obscure platforms

    Last time I checked, I got impressed by the wide platform support, once you go down the tier list (https://doc.rust-lang.org/nightly/rustc/platform-support.htm...). What "obscure platform" specifically are you thinking about, that is currently missing from those lists?

    • Just to be clear, I think this is a very small disadvantage. GCC and therefore C/C++ supports some old stuff like SuperH, Intel Itanium, PA-RISC and a bunch of microcontroller archs that LLVM does not.

      However, there is now a Rust codegen plugin for GCC, so even this disadvantage is now basically moot.

      Rewrite all the things!

      1 reply →

    • The list of supported GCC architectures is here: https://gcc.gnu.org/backends.html

      LLVM doesn't have an equivalent annotated list, but it is missing alpha, bfin, c6x, fr30, frv, gcn, h8300, ia64 (aka Itanium), lm32, m32c, m32r, mcore, mep, microblaze, mmix, mn10300, moxie, nds32, nios2, pa (aka PA-RISC), pdp11, pru, rl78, rs6000, rx, sh (aka SuperH), storm16, v850, vax, and visium. For its part, LLVM does have some targets that GCC doesn't have (mostly related to GPU compilation).

      The "big" targets that GCC has that LLVM lacks are Alpha, Itanium, PA-RISC, and SuperH, with Itanium being sufficiently weird that it's pretty firmly in the "fuck this" category from a maintainer's perspective, and people are actively ripping out support for it.

My very naive understanding is that a part of what makes mold fast is concurrency, which I'd expect to be a lot more error-prone in C/C++. Not having to worry about data races might give more confidence with trying out more complex techniques for how to split up work in a way that ends up making things faster

I suspect fearless concurrency is another motivating factor. Better perf can be achieved by squeezing more parallelism, but without borrow checking it's difficult to do fine-grained parallelism correctly.

  • Is "fearless concurrency" a technical term? I thought it was just the catchy name of a chapter of the Rust book.

It's possible to write safe code in C or C++ but it's extremely difficult to read existing code and prove it's safe, without using as much effort as it takes to write it in the first place. This includes the code you wrote last month whose surrounding code has changed. And you have to be right every time while the attacker only needs you to be wrong once. The problem is not writing the code, it is continually verifying it.

  • As someone who worked professionally with both C++ and Rust, I mostly agree with this, but I would say that writing the concurrent code in C++ is also hard.

    The only way I found that works reliably is stick to a small well defined set of mostly safe primitives. E.g. at work we use message passing / event bus architecture everywhere, which works great for a robotics / industrial context. But even then, if you somehow mess up and have a variable accessed from event handlers in two different threads, it is tough to spot other than if you get lucky and observe it with a build using TSAN.

    With rust that class of mistakes is just entirely eliminated, which makes it easier to to concentrate on the hard things that actually matter (like the domain specific logic).

> And second, did you use any AI tools for the rewrite?

IMO, given the recent commits: almost certainly.

> in a C/C++ code base

It's like "in a Zodiac boat / aircraft carrier navy". This customary putting C and C++ into the same bucket is as amusing as it is unproductive.

  • Apparently the old mold version had a couple of plain C source files in the source tree (both vendored 3rd party libs but also in the 'regular' source code), so technically C/C++ is correct in this case ...and looks like the Rust version also has some (very minimal) C code left (looks like mostly varargs stuff, I guess Rust doesn't have a concept of varargs?), so it could be called a C/Rust project ;)

    https://github.com/rui314/mold/tree/main/c