This paper presents the development and experimental validation of a reinforcement learning (RL)-based magnetic controller on the DIII-D tokamak. The controller directly maps raw magnetic diagnostic signals to actuator commands, replacing the traditional isoflux control algorithm based on equilibrium reconstruction. Four RL controllers are trained using the Soft Actor–Critic algorithm with an asymmetric Actor–Critic architecture in the NSFsim simulator. All controllers are deployed in the DIII-D Plasma Control System and operated with a 4 kHz feedback loop. Two randomization strategies are evaluated during training: evolving kinetic profiles and fixed kinetic profiles within each episode. The latter approach is found to better capture experimental deviations in the current density profile and to provide overall improved control performance. Robust operation is demonstrated across heating power scans in both L- and H-mode plasmas, as well as during transient events such as L–H transitions and pellet injections. Control errors in plasma shape and radial position remained within 1.5–2.0 cm and 1 cm, respectively. A notable discrepancy was observed in the vertical X-point position, with errors of up to approximately 4 cm, attributed to the current density distribution mismatches between simulations and experiments.
本論文は、DIII-Dトカマクにおける強化学習(RL)ベースの磁気制御器の開発と実験的検証を提示する。この制御器は、生の磁気診断信号をアクチュエータ指令に直接マッピングし、平衡再構成に基づく従来のアイソフラックス制御アルゴリズムを置き換える。4つのRL制御器は、NSFsimシミュレータにおいて非対称Actor–Criticアーキテクチャを用いたSoft Actor–Criticアルゴリズムによって訓練される。すべての制御器はDIII-Dプラズマ制御システムに実装され、4 kHzのフィードバックループで動作する。訓練中に2つのランダム化戦略が評価される: 各エピソード内で時間発展する運動論的プロファイルと、各エピソード内で固定された運動論的プロファイル。後者のアプローチは、電流密度プロファイルの実験的偏差をよりよく捉え、全体的に改善された制御性能を提供することが見いだされた。LモードおよびHモードプラズマの両方における加熱パワースキャン、ならびにL–H遷移やペレット入射などの過渡事象中において、ロバストな動作が実証される。プラズマ形状および半径方向位置の制御誤差は、それぞれ1.5〜2.0 cmおよび1 cm以内に留まった。垂直X点位置には顕著な不一致が観察され、誤差は約4 cmに達したが、これはシミュレーションと実験間の電流密度分布の不一致に起因する。