Comment by armcat
20 hours ago
The real story here is this wonderful exposition in applying diffusion models to a time series data that is neither discrete nor continuous. It’s always fascinating to see diffusion models applied in different scenarios, same with diffusion language models.
I don’t understand how anyone can say market data is continuous. It’s an aggregate of discrete orders and transactions. It may look continuous if you squint, but it’s not.
(Yes I read the article)
The article mentioned this at the beginning and end:
"But what kind of data is market data? ... is it continuous or discrete? Market data seems to have features of both. An order book evolves via a series of turns as market participants put on, or take away, resting orders; there is no doubt a discrete action space. And yet, many of the most important parameters of a given order, like its price, have such a high cardinality that they basically appear continuous. ...
One way to explore these complications is to treat your data as if it’s continuous and see how that breaks down. "
"Most of the project was spent grappling with a fundamental problem: market data is neither fully continuous nor fully discrete."
I think it is continuous because it is an emergent property of the market system. There is a video where the guy tried to see if market structure was "real" by modeling it. He had trouble in the beginning but when he added an order book the structures and patterns (Head and Shoulder, Bull Flags, Double Tops) appeared. It's called:
"I made a Market Simulation to see if Patterns are Real" by Krafer
https://www.youtube.com/watch?v=oWheof70O9g
What I liked about it is that they played around with different categories so as to tackle the natural discontinuity of how markets work/happen. With different categories the applicable models and data conditioning change. I think the requirement of manual tuning of the data and categories is touching bitter lesson aspects: a more general model would train and find the categories in its latent space.
I think the pattern In these posts is that they are looking for models that are explainable and have the potential for low latency, like the kann fpga project. While more general and opaque classifiers possibly work they are unlikely to work at the speed and risk constraint required to make money.
I apologise if I completely missed the point, it is not my field at all.
I think you’re the only other person that read the article.
Maybe other people did, they just chose to talk about something slightly off topic.
There is nothing in the community guidelines about having to stay on topic and to only talk about what's on the article.
There is however:
https://news.ycombinator.com/newsguidelines.html
What makes you say that?
Everyone else just comments and assumes based on the headline