The real story here is this wonderful exposition in applying diffusion models to a time series data that is neither discrete nor continuous. It’s always fascinating to see diffusion models applied in different scenarios, same with diffusion language models.
I don’t understand how anyone can say market data is continuous. It’s an aggregate of discrete orders and transactions. It may look continuous if you squint, but it’s not.
What I liked about it is that they played around with different categories so as to tackle the natural discontinuity of how markets work/happen. With different categories the applicable models and data conditioning change.
I think the requirement of manual tuning of the data and categories is touching bitter lesson aspects: a more general model would train and find the categories in its latent space.
I think the pattern In these posts is that they are looking for models that are explainable and have the potential for low latency, like the kann fpga project. While more general and opaque classifiers possibly work they are unlikely to work at the speed and risk constraint required to make money.
I apologise if I completely missed the point, it is not my field at all.
Maybe other people did, they just chose to talk about something slightly off topic.
There is nothing in the community guidelines about having to stay on topic and to only talk about what's on the article.
There is however:
Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
There is no model of the market that can remain stably accurate because the market will inevitably incorporate the insights of any model that is accurate until those insights are no longer accurate
Also known as the trading model's shelf-life. Based on what I have heard and learned, the average active trading model has a useful life of around 18 months. After that the rest of the market has adjusted and its edge is gone.
If rumours are to be believed, several "slow" hedge funds have models that remain useful and profitable for 5 years or more. Then again, those models are not used to conduct exchange trades but rely more on aspects of fundamental analysis.
Rather curiously some of the largest banks and tech-heavy asset managers tend to sponsor meetups and events like PyData - every now and then doing a slot on some of their open-sourced stuff. You might see a talk on 5-year-old trading model internals with large chunks of it opened up, or they could present some of the internal UI or visualisation libraries. The latter tend to be things that they no longer actively develop (they're "ready") but maintain for their ongoing persistent needs.
Of course anything that actually brings them money and/or gives an edge is not even discussed.
in order for your model to accomplish that, you would get very rich.
there will also likely always be more. in the limit in order to get an edge your model would start to infer insider information. for example, it's common knowledge by now that satellite imagery is used to measure car numbers in parking lots, that's a proxy for insider information.
so it's not even so much the model as it is the data.
even being able to forecast weather better than publicly available methods can be leveraged to gain a significant edge.
Just having a model that's predictably biased the wrong way when other models have already priced in the same things is pretty good. I mean, you have to look at where your model has no edge and figure out if it does have edges that the other ones are missing.
Coming at this from baseball, where the art is really knowing when not to take a game based on the line; usually if what you predict lines up too closely with Vegas, your edge is gone. (Hand-written A-life/evolutionary algorithm model of competing baseball equations that I've been refining for years, picks lately filtered through AI to isolate which ones fall into bands that are worth risking money on).
Remarkable internship. Fairly dense write-up too. Everything is informative.
Amusing degree of detail, though it makes sense it’s targeted at future interns. Why would you expect order inter-arrival to be normal? Surely an order arriving sort of boosts the probability of others, Hawkes-like. A nice little trick to get students talking I suppose.
There is a great line in the Man Who Solved The Market [0] about RenTech:
Paraphrasing:
"They went out of their way to try and fill in gaps in the historical market data. Over time, they got so good at predicting prices movements that they built algorithms to fill in the gaps."
Highly recommend the book if you are interested in the history of algorithmic trading.
Loved that book! Makes me wonder .. with how good AI is getting at solving novel math problems .. could you essentially bruteforce ur way towards some of RenTech’s techniques? Like .. have AI go through research papers (and related work) published by RenTech staff, figure out how their respective fields of expertise could be applied to trading, and run experiments/backtests to see what sticks? or is the hard part less the math and more the data and execution?
I have done a lot of research on RenTech. AFAIK, their early stuff isn't that much of a secret. Lots of Markov modeling and NLP-related stuff... but ultimately I believe it just came down to being one of the few firms actually even attempting statistical modeling with market data. (You had Ed Thorp too!)
You can easily find the backgrounds of people that were hired/worked there. Also, some amount of leakage over the years, bits and pieces, little clues... Deriving an approximation of "this was their strategy" with any amount of precision would almost be impossible. There is probably only a handful of people in the world who could tell you if you're hot or cold.
Very interesting thought experiment though...
> the hard part less the math and more the data and execution?
Ironically, this _was_ the part they had some trouble getting right. Programming was not nearly as ubiquitous back then.
For images, latent diffusion only works when using a special decoder that can produce realistic images given relevant samples from the latent space. This is trained like a GAN and doesn’t focus on pixel level error but higher level image features extracted via another network etc. expect that building a decoder like that for their problem may have solved their issues.
Plans don't determine outcome though, only broad sector movements eg announcing you are doing AI tend to produce a small spike in price, effect is slowly diminishing though.
The market has modes and reverts behavior when it switches them. Thus happy bouncy becomes hammered stammered. The prediction models fall hook and sinker for that.
For more info on why flow-matching is more stable than DDIM, see Heitz 2023, which nicely explains how they're almost equivalent but DDIM corresponds to a differential equation with a 1/α term that diverges at the start of the denoising process, while flow-matching/IADB doesn't diverge: https://arxiv.org/abs/2305.03486
> There were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better.
When they say "theoretically ideal" they mean a specific thing: the DDPM loss is information theoretically the number of error correction bits needed to recover an image. At very low noises, this quantity goes singular. This is not actually a problem—you should bound the log-SNR of your schedule by the image quantization level (usually 9 bits) anyway, and a trained bound lands within 10% of this.
Flow matching does not diverge because it removes the log-SNR rate from the loss. Since the rate can span several orders of magnitude over your schedule, you should importance sample during training. My guess is Kavish forgot to do this, leading to divergence.
This was really interesting and a lot to get done in a short internship, nevermind the public retrospective and hackernews appearance. Well done Kavish!
You hear that often, but if you squint your eyes, the entire idea of index funds is just that: they outperformed stock-pickers in the past, so you should put money into them to get higher returns in the future. There's no fundamental index fund investment thesis other than "past performance is indicative of future returns".
That thesis is at least to some extent self-fulfilling, because there's so much money flowing into index funds that prices of all the underlying assets keep moving up, and there's probably not enough money trying to bid against that / arbitrage the excesses away.
A similar thing could happen with AI. Markets are efficient only if the world isn't in some sort of a trance.
That's not the idea behind index funds. It's arithmetic. The aggregate return of active investors, before fees, is the market return. Once you subtract fees, it's below the market return. While some active managers' performance less fees is higher than the market return, it's very difficult to predict which will perform this way. So your best bet is to own the market through a broad index fund that has almost no cost.
If you want to read about this, see Sharpe (1991), The Arithmetic of Active Management.
No, the investment thesis is that there is a risk premia associated with investing in equity (as opposed to cash) that you will be compensated for - if this risk premia exists the way to harvest it is to be as diversified as possible (hand waving).
The relative performance of different baskets of equity is of course much more subtle / prone to behavioural effects and so on.
> There's no fundamental index fund investment thesis other than "past performance is indicative of future returns".
(Not a finance guy) But isn’t it that buy and hold index funds minimise trading fees, and those savings compound to produce better long term performance than nearly all active funds?
(The “Acquired” podcast episode covering the history of Vanguard and Jack Bogle goes into it in detail)
I don't think that's a particularly accurate assessment of the idea behind index funds.
The point of index funds isn't "these outperform all pickers, so they'll outperform all pickers in the future."
I think the idea is more around a combination of:
- you'll have much lower risk trying not to pick the right picker (or pick the investments yourself)
- the median picker is probably not very good (approached in two directions: sizable pickers that hit on an edge will likely be copied until the edge is gone, and smaller pickers are extremely unlikely to have enough specialized info or skills to excel).
> you should put money into them to get higher returns in the future.
Higher returns than what? I thought the whole point of buying broad market index funds was to simply get the market returns. For this thesis to make sense, you simply must assume that companies, in aggregate, make money - not that any particular company will follow past performance. If you don't think companies make money, then what are you doing buying equities?
Might not work for long holds but for short daytrading I feel like getting enough data for a model, not llm but any model, is the real golden goose moat of most ibs. Pajama traders at home have to set up heuristics for what they believe is a bull flag and maybe develop even a refined gut sense of spotting say a bull flag.
But imagine a quant at Jane street. They see the same candlestick pattern as the pajama guy but their model is giving them actual odds ratios instead of gut instinct. They can now score their putative bull flags in real time and make investments that might be more likely to pay off than not.
A big reason why this works is that technical analysis is a self fulfilling prophecy. Many people are looking for and trading on the exact same signals and this is enough to see a pattern in the candlestick data actually be one associated with market movement. Whether the market movement is 'genuine' or manufactured by other quantitative technical traders in this self fulfilling prophecy doesn't matter, you've made your money and really don't care about the underlying asset at the end of the day, only its delta.
WIthin markets the idea "it implies correlation" may be enough to shift your odds from 50/50 to 60/40, which is enough that people with "implying only" are usually fine to dare to enter a position.
The real story here is this wonderful exposition in applying diffusion models to a time series data that is neither discrete nor continuous. It’s always fascinating to see diffusion models applied in different scenarios, same with diffusion language models.
I don’t understand how anyone can say market data is continuous. It’s an aggregate of discrete orders and transactions. It may look continuous if you squint, but it’s not.
(Yes I read the article)
What I liked about it is that they played around with different categories so as to tackle the natural discontinuity of how markets work/happen. With different categories the applicable models and data conditioning change. I think the requirement of manual tuning of the data and categories is touching bitter lesson aspects: a more general model would train and find the categories in its latent space.
I think the pattern In these posts is that they are looking for models that are explainable and have the potential for low latency, like the kann fpga project. While more general and opaque classifiers possibly work they are unlikely to work at the speed and risk constraint required to make money.
I apologise if I completely missed the point, it is not my field at all.
I think you’re the only other person that read the article.
Maybe other people did, they just chose to talk about something slightly off topic.
There is nothing in the community guidelines about having to stay on topic and to only talk about what's on the article.
There is however:
https://news.ycombinator.com/newsguidelines.html
What makes you say that?
1 reply →
There is no model of the market that can remain stably accurate because the market will inevitably incorporate the insights of any model that is accurate until those insights are no longer accurate
Also known as the trading model's shelf-life. Based on what I have heard and learned, the average active trading model has a useful life of around 18 months. After that the rest of the market has adjusted and its edge is gone.
If rumours are to be believed, several "slow" hedge funds have models that remain useful and profitable for 5 years or more. Then again, those models are not used to conduct exchange trades but rely more on aspects of fundamental analysis.
Rather curiously some of the largest banks and tech-heavy asset managers tend to sponsor meetups and events like PyData - every now and then doing a slot on some of their open-sourced stuff. You might see a talk on 5-year-old trading model internals with large chunks of it opened up, or they could present some of the internal UI or visualisation libraries. The latter tend to be things that they no longer actively develop (they're "ready") but maintain for their ongoing persistent needs.
Of course anything that actually brings them money and/or gives an edge is not even discussed.
I would imagine a fundamental analysis model should remain fairly consistent over the years. It should be immune to a red queen situation.
1 reply →
This assumes that any accurate model will inevitably get big enough to be noticable by the rest of the actors.
in order for your model to accomplish that, you would get very rich.
there will also likely always be more. in the limit in order to get an edge your model would start to infer insider information. for example, it's common knowledge by now that satellite imagery is used to measure car numbers in parking lots, that's a proxy for insider information.
so it's not even so much the model as it is the data.
even being able to forecast weather better than publicly available methods can be leveraged to gain a significant edge.
> that's a proxy for insider information
Niche information pieces like this are only relevant to a veeeery small subset of all market participants, nearly invisible.
Just having a model that's predictably biased the wrong way when other models have already priced in the same things is pretty good. I mean, you have to look at where your model has no edge and figure out if it does have edges that the other ones are missing.
Coming at this from baseball, where the art is really knowing when not to take a game based on the line; usually if what you predict lines up too closely with Vegas, your edge is gone. (Hand-written A-life/evolutionary algorithm model of competing baseball equations that I've been refining for years, picks lately filtered through AI to isolate which ones fall into bands that are worth risking money on).
AKA the efficient markets hypotheses.
Or, thankfully, for Jane Street: the (eventually) efficient market hypothesis
There's definitely alpha out there, but I wouldn't want to make it my job to look for it
1 reply →
Remarkable internship. Fairly dense write-up too. Everything is informative.
Amusing degree of detail, though it makes sense it’s targeted at future interns. Why would you expect order inter-arrival to be normal? Surely an order arriving sort of boosts the probability of others, Hawkes-like. A nice little trick to get students talking I suppose.
Jane Street interns impressive as always.
There is a great line in the Man Who Solved The Market [0] about RenTech:
Paraphrasing:
"They went out of their way to try and fill in gaps in the historical market data. Over time, they got so good at predicting prices movements that they built algorithms to fill in the gaps."
Highly recommend the book if you are interested in the history of algorithmic trading.
0 - https://amzn.to/4jNFpOS
0 -
Just read that part recently, highly recommend the book too!
Loved that book! Makes me wonder .. with how good AI is getting at solving novel math problems .. could you essentially bruteforce ur way towards some of RenTech’s techniques? Like .. have AI go through research papers (and related work) published by RenTech staff, figure out how their respective fields of expertise could be applied to trading, and run experiments/backtests to see what sticks? or is the hard part less the math and more the data and execution?
I have done a lot of research on RenTech. AFAIK, their early stuff isn't that much of a secret. Lots of Markov modeling and NLP-related stuff... but ultimately I believe it just came down to being one of the few firms actually even attempting statistical modeling with market data. (You had Ed Thorp too!)
You can easily find the backgrounds of people that were hired/worked there. Also, some amount of leakage over the years, bits and pieces, little clues... Deriving an approximation of "this was their strategy" with any amount of precision would almost be impossible. There is probably only a handful of people in the world who could tell you if you're hot or cold.
Very interesting thought experiment though...
> the hard part less the math and more the data and execution?
Ironically, this _was_ the part they had some trouble getting right. Programming was not nearly as ubiquitous back then.
For images, latent diffusion only works when using a special decoder that can produce realistic images given relevant samples from the latent space. This is trained like a GAN and doesn’t focus on pixel level error but higher level image features extracted via another network etc. expect that building a decoder like that for their problem may have solved their issues.
Can you deduce a companies plans from its actions (hiring, ordering,) long before it makes them public via AI?
Can you find vc appetite for a startup before they buy equity
Plans don't determine outcome though, only broad sector movements eg announcing you are doing AI tend to produce a small spike in price, effect is slowly diminishing though.
The market has modes and reverts behavior when it switches them. Thus happy bouncy becomes hammered stammered. The prediction models fall hook and sinker for that.
For more info on why flow-matching is more stable than DDIM, see Heitz 2023, which nicely explains how they're almost equivalent but DDIM corresponds to a differential equation with a 1/α term that diverges at the start of the denoising process, while flow-matching/IADB doesn't diverge: https://arxiv.org/abs/2305.03486
You are referencing this from the article:
> There were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better.
When they say "theoretically ideal" they mean a specific thing: the DDPM loss is information theoretically the number of error correction bits needed to recover an image. At very low noises, this quantity goes singular. This is not actually a problem—you should bound the log-SNR of your schedule by the image quantization level (usually 9 bits) anyway, and a trained bound lands within 10% of this.
Flow matching does not diverge because it removes the log-SNR rate from the loss. Since the rate can span several orders of magnitude over your schedule, you should importance sample during training. My guess is Kavish forgot to do this, leading to divergence.
This was really interesting and a lot to get done in a short internship, nevermind the public retrospective and hackernews appearance. Well done Kavish!
No
Seconded.
"Past performance is not indicative of future results."
You hear that often, but if you squint your eyes, the entire idea of index funds is just that: they outperformed stock-pickers in the past, so you should put money into them to get higher returns in the future. There's no fundamental index fund investment thesis other than "past performance is indicative of future returns".
That thesis is at least to some extent self-fulfilling, because there's so much money flowing into index funds that prices of all the underlying assets keep moving up, and there's probably not enough money trying to bid against that / arbitrage the excesses away.
A similar thing could happen with AI. Markets are efficient only if the world isn't in some sort of a trance.
That's not the idea behind index funds. It's arithmetic. The aggregate return of active investors, before fees, is the market return. Once you subtract fees, it's below the market return. While some active managers' performance less fees is higher than the market return, it's very difficult to predict which will perform this way. So your best bet is to own the market through a broad index fund that has almost no cost.
If you want to read about this, see Sharpe (1991), The Arithmetic of Active Management.
5 replies →
No, the investment thesis is that there is a risk premia associated with investing in equity (as opposed to cash) that you will be compensated for - if this risk premia exists the way to harvest it is to be as diversified as possible (hand waving).
The relative performance of different baskets of equity is of course much more subtle / prone to behavioural effects and so on.
> There's no fundamental index fund investment thesis other than "past performance is indicative of future returns".
(Not a finance guy) But isn’t it that buy and hold index funds minimise trading fees, and those savings compound to produce better long term performance than nearly all active funds?
(The “Acquired” podcast episode covering the history of Vanguard and Jack Bogle goes into it in detail)
I don't think that's a particularly accurate assessment of the idea behind index funds.
The point of index funds isn't "these outperform all pickers, so they'll outperform all pickers in the future."
I think the idea is more around a combination of:
- you'll have much lower risk trying not to pick the right picker (or pick the investments yourself)
- the median picker is probably not very good (approached in two directions: sizable pickers that hit on an edge will likely be copied until the edge is gone, and smaller pickers are extremely unlikely to have enough specialized info or skills to excel).
> you should put money into them to get higher returns in the future.
Higher returns than what? I thought the whole point of buying broad market index funds was to simply get the market returns. For this thesis to make sense, you simply must assume that companies, in aggregate, make money - not that any particular company will follow past performance. If you don't think companies make money, then what are you doing buying equities?
The stocks in the index funds are picked, just by a committee. And there are a few competing index funds.
1 reply →
Might not work for long holds but for short daytrading I feel like getting enough data for a model, not llm but any model, is the real golden goose moat of most ibs. Pajama traders at home have to set up heuristics for what they believe is a bull flag and maybe develop even a refined gut sense of spotting say a bull flag.
But imagine a quant at Jane street. They see the same candlestick pattern as the pajama guy but their model is giving them actual odds ratios instead of gut instinct. They can now score their putative bull flags in real time and make investments that might be more likely to pay off than not.
A big reason why this works is that technical analysis is a self fulfilling prophecy. Many people are looking for and trading on the exact same signals and this is enough to see a pattern in the candlestick data actually be one associated with market movement. Whether the market movement is 'genuine' or manufactured by other quantitative technical traders in this self fulfilling prophecy doesn't matter, you've made your money and really don't care about the underlying asset at the end of the day, only its delta.
Answering this aphorism with a demonstration of your wrongheaded idea about timing the market is embarrassing stuff.
2 replies →
It is though, it just doesn't guarantee it.
Same as "correlation doesn't imply causation" - it actually does imply it, it just doesn't prove it.
WIthin markets the idea "it implies correlation" may be enough to shift your odds from 50/50 to 60/40, which is enough that people with "implying only" are usually fine to dare to enter a position.
[flagged]
[flagged]
[flagged]