This work develops an artificially intelligent (AI) tokamak operation design algorithm that provides an adequate operation trajectory to control multiple plasma parameters simultaneously into different targets. An AI is trained with the reinforcement learning technique in the data-driven tokamak simulator, searching for the best action policy to get a higher reward. By setting the reward function to increase as the achieved βp, q95, and li are close to the given target values, the AI tries to properly determine the plasma current and boundary shape to reach the given targets. After training the AI with various targets and conditions in the simulation environment, we demonstrated that we could successfully achieve the target plasma states with the AI-designed operation trajectory in a real KSTAR experiment. The developed algorithm would replace the human task of searching for an operation setting for given objectives, provide clues for developing advanced operation scenarios, and serve as a basis for the autonomous operation of a fusion reactor.
本論文は、複数のプラズマパラメータを同時に異なる目標値へ制御するための適切な運転軌道を提供する、人工知能(AI)を用いたトカマク運転設計アルゴリズムを開発するものである。データ駆動型のトカマクシミュレータにおいて強化学習手法を用いてAIを訓練し、より高い報酬を得るための最適な行動方針を探索する。報酬関数を、達成されたβp、q95、liがそれぞれの目標値に近づくほど増加するように設定することで、AIは目標値を満たすようにプラズマ電流と境界形状を適切に決定することを試みる。様々な目標値と条件下でシミュレーション環境においてAIを訓練した後、実機のKSTAR実験において、AIが設計した運転軌道によって目標とするプラズマ状態を達成できることを実証した。本アルゴリズムは、与えられた目標に対する運転設定値の探索という人間の作業を代替し、高度な運転シナリオの開発のための指針を提供するとともに、核融合炉の自律運転の基盤となるものである。