Comment by miellaby
8 hours ago
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
How much should German authors be paid, $3,000 per book, like Anthropic paid?
Why expose yourself to this liability?
How do you cleanse the data at this scale?
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
4 replies →
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
How, is the training data public?
Isnt OLMO basically open like this? My understanding is people have recreated the model from the same data sources with repeatable responses, or reasonably close to the original (since LLMs never answer the same).
LLMs can be configured to basically be deterministic. It's only really bad for performance because language is not deterministic.
The paper mentioned is here: https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)
I wish universities would take it upon themselves to curate the training sets for these models.
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
Yeah the pdf alone is awesome as a learning tool.
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.
Bravo team!
Hopefully this becomes the new standard.
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
Be cautious what you wish for. Tools don't tell you what to do with them.
Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.
Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.
Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.
Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
7 replies →
The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.
But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.
As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?
7 replies →
Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.
Wonder what safeguards Kolibri uses to prevent this? Or if they even can
8 replies →
Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.
But the alternative just seems… so much worse to me?
A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.
And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.
3 replies →
I think the unethical things are happening already with boutique firms. I admit I don't get the concern.
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.
As is shown to us by filthy rich people every day.
Or did you mean the burglar in the fawellas?
[dead]
His point about regulation and innovation is great and I wish more people thought like that.
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)