← Back to context

Comment by _wire_

1 year ago

Solid overview of applied color theory for video, so worth watching.

As to what was to be debunked, the presentation not only fails to set out a thesis in the introduction, it doesn't even beg a question, so you've got to watch hours to get to the point: SDR and HDR are two measurement systems which when correctly used for most cases (legacy and conventional content) must produce the visual result. The increased fidelity of HDR makes it possible to expand the sensory response and achieve some very realistic new looks that were impossible with SDR, but the significance and value of any look is still up to the creativity of the photographer.

This point could be more easily conveyed by this presentation if the author explained in the history of reproduction technology, human visual adaptation exposes a moment by moment contrast window of about 100:1, which is constantly adjusting across time based on average luminance to create an much larger window of perception of billions:1(+) that allows us to operate under the luminance conditions on earth. But until recently, we haven't expected electronic display media to be used in every condition on earth and even if it can work, you don't pick everywhere as your reference environment for system alignment.

(+)Regarding difference between numbers such as 100 or billions, don't let your common sense about big or small values phase your thinking about differences: perception is logarithmic; it's the degree of ratios that matter more than the absolute magnitude of the numbers. As a famous acoustics engineer (Paul Klipsch) said about where to focus design optimization of response traits of reproduction systems: "If you can't double it or halve it, don't worry about it."

It's hard to boil it down to a simple thesis because the problem is complicated. He admits this in the presentation and points to it being part of the problem itself; there are so many technical details that have been met with marketing confusion and misunderstanding that it's almost impossible to adequately explain the problem in a concise way. Here's my takeaway:

- It was clearly a mistake to define HDR transfer functions using absolute luminance values. That mistake has created a cascade of additional problems

- HDR is not what it was marketed to be: it's not superior in many of the ways people think it is, and in some ways (like efficiency) it's actually worse than SDR

- The fundamental problems with HDR formats have resulted in more problems: proprietary formats like Dolby Vision attempting to patch over some of the issues (while being more closed and expensive, yet failing to fully solve the problem), consumer devices that are forced to render things worse than they might be in SDR due to the fact that it's literally impossible to implement the spec 100% (they have to make assumptions that can be very wrong), endless issues with format conversions leading to inaccurate color representation and/or color banding, and lower quality streaming at given bit rates due to HDR's reliance on higher bit depths to achieve the same tonal gradation as SDR

- Not only is this a problem for content delivery, but it's also challenging in the content creation phase as filmmakers and studios sometimes misunderstand the technology, changing their process for HDR in a way that makes the situation worse

Being somewhat of a film nerd myself and dealing with a lot of this first-hand, I completely agree with the overall sentiment and really hope it can get sorted out in the future with a more pragmatic solution that gives filmmakers the freedom to use modern displays more effectively, while not pretending that they should have control over things like the absolute brightness of a person's TV (when they have no idea what environment it might be in).

  • While HDR has the problems described by you, in practice, whenever possible, I choose the HDR version of a movie over its SDR version.

    The reason is not HDR itself, but the fact that the HDR movies normally use the BT.2020 color space, while the SDR movies normally use the BT.709 color space.

    The color spaces based on the limitations of the first color CRT tubes, which are no longer relevant today, i.e. sRGB, BT.709 and the like, are really unacceptable from my point of view, because they cannot reproduce many of the more saturated colors in the red-orange region, which are frequently encountered in nature and in manufactured objects, and which are also located in a region of the color space where human vision is most sensitive.

    While no cheap monitor can reproduce the full BT.2020, many cheap monitors can reproduce the full DCI-P3 color space, which provides adequate improvements in the red-orange corner over sRGB/BT.709.

    • I assume you've watched the "Wider Gamut Misinformation" portion of the video. You're correct in saying that HDR normally uses Rec. 2020 (because Rec. 2100 points to Rec. 2020's color primaries), but Steve points out that Rec. 2020 is an SDR color space which doesn't technically require HDR.

      It's true that, given the options available today, the two usually go hand-in-hand (wider color gamut and HDR). However, one of the arguments Steve makes is that a majority of content (and a vast majority of the pixels in that content) doesn't use colors outside Rec. 1886's gamut. Illuminated objects (natural or manmade) almost never go outside that, so you're usually only talking about a few pixels from intensely-saturated light sources in the shot (like LEDs) that might use those colors. Even then, not a lot of filmmakers feel the need to go there, so their movies will look the same in narrow and wide gamuts.

      I don't think the video is arguing against wider gamut or even higher dynamic range as options; modern displays are more capable than older ones, so we need tools to allow content creators to use that capability if they desire. The problem is that all of these things (color space, bit depth, transfer function, absolute luminance values, etc) have been lumped together under one label, "HDR", and some of the implementation details are actually worse than what we had with SDR. If you skip to the "Checklist Recap" portion of the video, you'll see that there are actually quite a few downsides to HDR in its current form, but since most of the standards are tightly coupled, we're kind of stuck unless we move to something better.

      I also personally choose HDR versions when watching movies, but that's because UHD content is usually also HDR. What I really want is the higher resolution. I've never felt like I'd be missing out if it didn't have HDR because I've compared the two a lot - they're really mostly the same for the movies I watch with a properly calibrated screen. To each their own :)

      3 replies →

Regardless of whether it is HDR or SDR, when processing raw data for display spaces one must throw out 90%+ of information of what was captured by the sensor (which is often a small amount of what was available at the scene already). There can simply be no objectivity, it is always about what you saw and what you want others to see, an inherently creative task.

  • 90% really ? What color information get ejected exactly ? For the sensor part are you talking about the fact that the photosites don't cover all the surface ? Or that we only capture a short band of wavelength ? Or that the lens only focuses rays unto specific exact points and make the rest blurry and we loose 3D ?

    • Cameras capture linear brightness data, proportional to the number of photons that hit each pixel. Human eyes (film cameras too) basically process the logarithm of brightness data. So one of the first things a digital camera can do to throw out a bunch of unneeded data is to take the log of the linear values it records, and save that to disk. You lose a bunch of fine gradations of lightness in the brightest parts of the image. But humans can't tell.

      Gamma encoding, which has been around since the earliest CRTs was a very basic solution to this fact. Nowadays it's silly for any high-dynamic image recording format to not encode data in a log format. Because it's so much more representative of human vision.

      5 replies →

    • Third blind man touching the elephant here: the other commenters are wrong! it’s not about bit depth or linear-to-gamma, it’s the fact that the human eye can detect way more “stops” (the word doesn’t make sense you have to just look it up) of brightness (I guess you could say “a wider range of brightness”, but photography people all say “stops”) than the camera, and the camera can detect more stops of brightness than current formats can properly represent!

      So you have to decide whether to lose the darker parts of the image or the brighter parts of the image you’re capturing. Either way, you’re losing information.

      (In reality we’re all kind of right)

      3 replies →

    • A 4k 30fps video sensor capturing 8 bits per pixel (bayer pattern) image, is capturing 2 gigabits per second. That same 4k 30fps video on Youtube will be 20 megabits per second or less.

      Luckily, it turns out relatively few people need to record random noise, so when we lower the data rate by 99% we get away with it.

      3 replies →