Many measurements are required to control thermonuclear plasmas and to fully exploit them scientifically. In the last years JET has shown the potential to generate about 50 GB of data per shot. These amounts of data require more sophisticated data analysis methodologies to perform correct inference and various techniques have been recently developed in this respect. The present paper covers a new methodology to extract mathematical models directly from the data without any a priori assumption about their expression. The approach, based on symbolic regression via genetic programming, is exemplified using the data of the International Tokamak Physics Activity database for the energy confinement time. The best obtained scaling laws are not in power law form and suggest a revisiting of the extrapolation to ITER. Indeed the best non-power law scalings predict confinement times in ITER approximately between 2 and 3 s. On the other hand, more comprehensive and better databases are required to fully profit from the power of these new methods and to discriminate between the hundreds of thousands of models that they can generate.
熱核プラズマを制御し、科学的に完全に活用するためには、多くの測定が必要である。近年、JETは1ショットあたり約50GBのデータを生成する可能性を示している。このような大量のデータには、より高度なデータ解析手法による正確な推論が必要であり、これに関して様々な技術が近年開発されている。本論文では、その表現形式に関する事前の仮定を置かずに、データから直接数学モデルを抽出する新しい手法について述べる。このアプローチは、遺伝的プログラミングによる記号回帰に基づいており、エネルギー閉じ込め時間に関する国際トカマク物理活動データベースのデータを用いて実証されている。得られた最良のスケーリング則はべき乗則の形式ではなく、ITERへの外挿の再検討を示唆している。実際、最良の非べき乗則スケーリングは、ITERにおける閉じ込め時間を約2秒から3秒と予測している。一方で、これらの新しい手法の利点を完全に活かし、それらが生成する数十万のモデルを識別するためには、より包括的で質の高いデータベースが必要である。