Teaching a network the post-Newtonian terms cuts training data by ten thousand times
A purely neural waveform model needs 10^8 to 10^10 training waveforms, but writing the known physics into the network and letting it learn only the corrections takes 8 numerical-relativity waveforms.
The argument · tap a timestamp to hear it
LIGO's error bars may come from theory, not the instrument
The speaker says that for some LIGO events, it cannot be ruled out that the analysis result is limited by understanding of waveform systematic errors, not just cryogenic cooling or mirror defects. GW23123 is an example: run different models and the posterior distributions show non-trivial differences. His position is not to take sides on who is right, but to give the models better error bars so the data analysts can keep arguing about who is right.
— Nils DeppeDouble the simulation length, quadruple the error
The mismatch of numerical-relativity simulations scales roughly as the square of the wavelength. Because simulation accuracy is dominated by phase error, phase error accumulates linearly in time, and mismatch is the square of phase error. So wanting a simulation twice as long does not just mean waiting longer, it means making the accuracy higher too, and the cost is penalised multiplicatively. Most simulations today land at about 30 orbits, while the reference line is about 25 orbits.
— Nils DeppeOn next-generation detectors, pure numerical relativity has no hope
The speaker showed a plot: the horizontal axis is event signal-to-noise ratio, the vertical axis is the number of simulations in the catalogue that are accurate and long enough. For LIGO, about 85% to 90% of the catalogue is good enough, roughly 3000 simulations; but for heavy systems at a signal-to-noise ratio of 10^4, only about 10 simulations in total are good enough. For Cosmic Explorer and Einstein Telescope, the conclusion is that numerical relativity alone is completely insufficient and other models are needed to fill the gap.
— Nils DeppeHybrid waveforms are actually stitching together two different physical systems
When splicing post-Newtonian waveforms onto numerical-relativity waveforms, the cost function has many local minima in parameter space, and the global minimum sits in a narrow, shallow valley that optimisers struggle to find. The test is to run a simulation of about 60 orbits and then move the matching window to see what the early waveform looks like. The result is that the early numerical-relativity waveform and the post-Newtonian waveform look completely different, which amounts to smoothly sewing together the waveforms of two different physical systems and introducing a bias.
— Nils DeppeLet the network learn only what post-Newtonian cannot compute
Each additional post-Newtonian term is a PhD thesis worth of work, so the speaker's idea is to treat the neural network as encoding all the post-Newtonian terms he does not want to write by hand. The network needs only about 100 parameters, about 120 CPU hours of training, and only 8 numerical-relativity waveforms in mass-ratio space. Besides the waveform difference, the loss function also requires the orbital velocity to be close, the correction term to vanish as the orbital velocity goes to zero, and equal-mass symmetry and weight regularisation.
— Nils DeppePhysics-informed networks extrapolate better outside the training set
At a mass ratio of 9 to 5, a point outside the training set, post-Newtonian completely fails near merger, the surrogate model (trained up to mass ratio 8) gets worse as the extrapolation error grows, while the physics-informed neural network does much better. If you then add a data point at mass ratio 15 after having trained up to mass ratio 8, the error all the way up to mass ratio 15 drops substantially. This promises to reduce the amount of training data needed to build a model.
— Nils DeppePurely neural approaches need more than ten thousand times the training data
The speaker says that most other attempts he has seen at modeling waveforms with neural networks throw physics away entirely and use only neural networks or autoencoders, and that really does need 10^8 to 10^10 waveforms to train, which means you would have to build a surrogate first in order to generate the training data, because you cannot get that many numerical-relativity waveforms. His conclusion: machine learning is a useful tool, but do not forget that you actually do know something about the thing you are modeling.
— Nils DeppeThe surrogate's 30 to 40 times speedup comes from changing the algorithm
The speaker says all he did was make the code less bad, with no revolutionary insight. The starting point was about 30 milliseconds, and now there is about a 30 to 40 times speedup. The biggest speedup came from changing the algorithm: from cubic interpolation to tenth-order Hermite polynomials, with far fewer time nodes. At present the dynamics part dominates, at about 370 microseconds, while waveform evaluation takes only 50 microseconds. He estimates that in the end it can deliver a frequency-domain waveform in 0.2 milliseconds, comparable to or better than the fastest phenom waveforms.
— Nils DeppeIn their own words · checked verbatim
anything that you agree with you should attribute to them anything that you disagree with you should attribute to me
Nils Deppe1:18
if I want a simulation that's twice as long in time, I will have a total error that's four times larger in a 3D simulation
Nils Deppe8:25
you're taking the PN wave from form from one system early on and stitching it to together with a waveform from a different physical system and then you combine them in a smooth way
Nils Deppe23:40
calculating post Newtonian terms is so difficult that like at this stage it's like one extra term is an entire PhD thesis
Nils Deppe33:50
you should remember that you have physics like machine learning is is is a useful tool but don't forget that you actually maybe know something about the thing you're modeling
Nils Deppe44:58
I haven't done roughly speaking anything more than like fired a few neurons right like none of this was like I would say revolutionary like insight deep things this is just picking the code and making it not suck
Nils Deppe55:08
I don't think I can give you a surrogate model that's faster than 0.2 two milliseconds for a frequency domain waveform for this kind of a system
Nils Deppe57:10
Figures
| Fraction of the LIGO catalogue that is good enough | about 85% to 90%, about 3000 simulations | 17:36 |
| Number of good-enough simulations for heavy systems at a signal-to-noise ratio of 10^4 | about 10 | 16:33 |
| Number of parameters in the physics-informed neural network | about 100 | 33:50 |
| Number of numerical-relativity waveforms needed to train the physics-informed neural network | 8 | 33:50 |
| Training time for the physics-informed neural network | about 120 CPU hours | 33:50 |
| Surrogate speedup factor | about 30 to 40 times | 55:08 |
| Waveform evaluation time | 50 microseconds | 54:08 |
| Dynamics part time | 370 microseconds | 55:08 |
| Target frequency-domain waveform time | 0.2 milliseconds | 57:10 |
| Target time for a single LIGO event analysis | under 30 minutes | 1:01:15 |
Glossary
- surrogate model
- A fast model that interpolates a waveform at arbitrary parameters from a small number of numerical-relativity simulations.
- physics informed neural network
- A network that has the known physics equations written into its structure and is left to learn only the correction terms.
- post-Newtonian
- An approximate analytic expansion for the early inspiral of a gravitational wave, less accurate the closer you get to merger.
- mismatch
- A metric for the difference between two waveforms, dominated by phase error in numerical relativity.
- co-orbital frame
- A frame that rotates with the binary, making the waveform smoother and easier to model.
How to listen
Graduate students and engineers working on gravitational-wave data analysis, numerical relativity or scientific machine learning, especially anyone who wants to use physical constraints to cut training data.
The first ~5 minutes of conference introductions and detector sensitivity curve background can be fast-forwarded.