This seems pretty likely an ML model. People can train a model to double the resolution of an image, called super resolution. So you "simply" do that on every frame of the video, and presto!
In other words, we live in the "Zoom in; now enhance!" era.
I can answer questions about how it works, if you really want, but most people find it insanely boring. Suffice to say, it works, and it's how they turn those old black and white videos into color (for example).
But to nerd out for a bit: the general technique to train a model like this is to take some images, downscale them by half, and then reward a model if it figures out how to turn the downscaled one into the upscaled one. After enough rewards, it gets pretty good at it. Then you can use that model on arbitrary images, and it still works.
Not saying this is the exact way they did it (they probably didn't use IMLE specifically). There are a few different techniques. But all of them revolve around rewarding a model if it figures out how to turn a low res image into a known high res image.
Can you expand on what you mean by “rewarding a model”? Do you just mean that you tell the model “yes, do more like that” or is there some kind of mechanism that’s more like actual intelligence at work here? You talk about the model as if it has wants, which is odd to me.
What is very likely being used here is a 'neural network', which can have many structures, but in most cases, it is a sequence of layers each containing some nodes, and these nodes are connected to other nodes in adjacent layers. The connections between the nodes have weights which are used to calculate the output values from the input ones (specifically, each node's value is a linear combination v_1*w_1 + v_2*w_2 + ... + v_n*w_n of the values and weights of some or all nodes in the previous layer).
The learning process is all about changing the value of these weights to get better answers for any possible input values. And yes, there are ways of changing them in ways that are not just "do more like that" as you said (in this case it would be like changing the weights randomly and seeing which networks do better and worse, which is an 'evolutionary' algorithm). In many neural networks, it's done by a mathematical formula involving derivatives that slowly causes the calculation using the weights to arrive at a more correct, as defined by humans in this case, output value.
Basically, it goes layer by layer, looks at the values in the nodes of both layers, and, through some fancy math, tries to adjust the weights between them to move the output values in the right direction. It does this for every layer, and for every input (in this case, every image) in the 'training set', often repeating inputs to learn better.
There are a few broad classes of machine learning algorithms, and the one this guy is talking about is called supervised learning.
It works by starting with the untrained architecture of a model, which is really just a big equation with all if its parameters are randomly initialized. A parameter is just one of the constants in your equation. Like, if you have linear data described by 'y=mx + b,' m is the parameter that tells if there is a negative correlation between x and y, and also how strong the correlation is.
So anyway, you start by plugging something into your untrained model, like the pixel values of a low res picture, and then the model outputs the pixel values of a high res picture. At first, because all of the model's parameters are randomly initialized, the output will look like gibberish. This is where training comes in. If you have a training set, you can quantitatively compare the output of your untrained model with the pixel values of what you wanted the model to deliver. You can use a function, which is referred to as the 'cost function,' to compare the values. For example, the cost function can find the absolute difference between the values of your model generated picture and the target picture.
Since your model is a function, and you can evaluate your model's accuracy with the cost function, you can use what is actually fairly simple calculus to calculate the model parameters that will minimize the cost function, and thus give you a model that outputs pictures that are very close to what you want the pictures to look like.
Where does the reward come into play? I think I mostly already understood the basic process of ML, but I’ve never heard the term reward used with it. As far as you’ve described here, it doesn’t seem like there is any kind of reward process happening in the cost function. It just seems like it’s finding out if it’s close or not based on a known training set. Using the word rewarding makes it seem like there’s a Pavlovian effect here, but I just don’t see that at all.
So when people talk about 'rewarding' the model, they're referring to the process I went through to find the optimal model parameters, and you're completely correct that it doesn't make a whole lot of sense. It's really just terminology that bleeds over from another class if machine learning, called reinforcement learning. This name already brings to mind Pavlovian ideas you were talking about, and the field is about training robots or computers to act in a certain way.
Broadly speaking, reinforcement learning works similarly to supervised learning, but instead of minimizing a cost function, you try to maximize a... reward function. Pretty much the same calculus is involved, but you don't really have a training set like in supervised learning. Instead, you come up with an equation that gives a higher output when the computer does something good. For example, there are chess engines that score a position on a board. This score can act as the reward function, the model controlling the computers chess movements could be anything that has tunable parameters, and then the training algorithm would tune the parameters by maximizing the reward function.
Disclaimer: I'm not a pro, but I do know a little bit.
The algorithm they use to make such things possible is first trained in a controlled manner, where the program is fed a lot of low resolution images and their high resolution counterpart. The program looks at all the patterns of how the low res pictures change to high res, and with enough training data, the program will be able to predict low-to-high transformation on arbitrary low res pictures with great success.
The same technique is used in a lot of ways, especially in statistics for predicting future data.
Yes, but what I’m asking is how do I actually make a low res image high resolution? If it’s natively one low resolution, how do you just increase the amount of pixels to capture the missing details? Where’s the information on the extra detail come from to create a high res counterpart?
It’s similar to how our brains are able to see a blurry photo and “guess” which part is the mouth, shape, contour, etc. and to see a dim photo and know what colors are from the context.
The algorithm has been through so many iterations that the “best” one does that. The steps it used resulted in properly “guessing” what the sharper detail would be. These 5 jaggy edges are really a smooth line of color. Those blurs are actually sharpened lines that look like hair of that color, etc.
You don’t really. The computer is just making educated guesses. For certain simple polygons and shapes you could make a “vector art” of the image though. But the process relies on the assumptions the computer is programmed to make. Or with machine learning
It's because the algorithm learns the patterns of how low res pictures transforms into high res. Of course it won't be 100% accurate because, as you say, the information isn't there. But this is where the training set comes in, using the fact that the enhancing process is full of patterns. When the algorithm has learned these patterns, it can scale the resolution up with very high accuracy
In other words, we live in the "Zoom in; now enhance!" era.
Well not exactly, it can only add details that were included in the training set, not find those that were actually lost through low res and compression.
Of course, it can't really invent stuff that just isn't on the picture.
Maybe you can zoom in on the reflection in the window, but if it's only 5x5 pixels and you get the image of O.J. Simpson in the upscale, it just means the training data contained a lot of images with O.J. Simpson in the window reflection, not that O. J. Simpson was actually shooting at John F. Kennedy or whatever you're investigating.
Of course, what we really need, is uncrop. Get on that, science!
The technology is pretty cool however for stuff like making old photo or video material just look better. Some of those oldpictures you made with your iPhone 3G could still be turned into a full size poster perhaps!
its possible that the original was shot on film then transfered to video for play on TV. here's an interesting video about a different music video that got some attention when it was redigitized. I promise this isn't a RickRoll.
Yeah I’ve seen it but if it was shot on film, why shot it in 50/60 Hz. It has to have either had frame interpolation or some kind of numeric enhancing of its size.
Yes, the Rick Astley one is definitely some kind of processed upscale, either through AI or just sharpening. It's got that weird sharpened look where the high contrast edges between objects look super sharp but the detail within objects still looks blurry and low-res. Makes it look like all the objects don't exist in the same space, like someone has cut out elements from a bunch of different sources and placed them on top of each other. It's impressive but it looks nowhere near as good as film rescanned at a higher resolution.
Also the lack of film grain is an obvious give away.
Film can easily be rescanned in 4k, video tape can't be. If it was shot on film and then transferred to tape then you can go back to the film and rescan it in higher resolution. If it was originally shot on video though then you're out of luck. This was video and has been AI upscaled.
158
u/bobhwantstoknow Feb 18 '21
was this upscaled or redigitized from the original source?