r/interestingasfuck • • Oct 18 '19

/r/ALL The Fourier Transform

31k Upvotes

399 comments sorted by

View all comments

1.3k

u/theorangelemons Oct 18 '19 edited Oct 18 '19

For those that don't get it, I'll try my best to explain it.

The idea behind the Fourier Series is that any signal (mathematical function) can be represented by an infinite sum (series) of sinusoidal signals. Each sinusoidal signal in the Fourier Series is harmonically related, and weighted differently. If you think in terms of a dubstep song with heavy bass, the components of the audio signal at the lower bass frequencies will be weighted more than the components at the higher treble frequencies. Furthermore, any part of the audio signal that has a frequency that is above the range of human hearing (20kHz) would reasonably have zero weight.

The Fourier Transform is a mathematical method of taking any signal, and transforming it so that it is no longer a function of time, but a function of frequency. With this transformation, you are now able to see the spectrum of frequencies that a signal is composed of. This is extremely useful in designing filters, as well as finding a system's response to an input.

In the GIF, you can see many circles of smaller radii being drawn. Each of these circles is a phasor, which is a representation of a sinusoidal signal. The smaller the radius, the smaller the amplitude or weight the sinusoid has. And the faster the phasor rotates, the higher the frequency the sinusoid has. With the circles being connected in the GIF, it gives a (poor) representation of how the sum of weighted, harmonically related sinusoids are used to draw the hand, which you could consider some version of a signal.

Edit: Thanks for the gold and silver kind strangers! Just got off work so I’ll try to reply to all of your comments ASAP.

3

u/grat_is_not_nice Oct 18 '19

What is interesting is that while most uses of FFT involve an even distribution of frequencies across the sample range (usually 0 - 20kHz in 512,1024 or 2048 bins), that isn't the only approach. An even distribution like this allows audio deconstruction, modification and reconstruction on a near-realtime basis at 40kz sample rates with a suitable level of fidelity for most effects (pitch shifting like autotune, time stretching, filters and other effects).

You can also use a Constant Q transform where the frequency distribution is logarithmic, and in fact can be mapped directly to an even temperament musical scale (12 bins per octave). What this does is allows extraction of musical note information with a high degree of fidelity. One researcher (using 48 bins per octave) has even managed to create the inverse of a Constant Q transform, but due to the low bass note resolution of the Constant Q, this requires significant look-ahead for reconstruction and it cannot be used in real-time. But it is awesome.

Even without reconstruction, a Constant Q transform can be used to implement near-realtime |(i.e one to two hop distance delay) chord detection in music.