논문 상세보기

Reinforcement Learning-Based Multi-Objective Arctic Route Optimization Considering Sea Ice Risk, Time, and Cost KCI 등재

해빙 위험도와 시간ㆍ비용을 고려한 강화학습 기반 북극항로 다목적 최적화

Seungwoo Lee, Byungheon Lee, Sunghyun Sim
  • 언어KOR
  • URLhttps://db.koreascholar.com/Article/Detail/452101
구독 기관 인증 시 무료 이용이 가능합니다. 4,200원
한국산업경영시스템학회지 (Journal of Society of Korea Industrial and Systems Engineering)
한국산업경영시스템학회 (Society of Korea Industrial and Systems Engineering)
초록

The Northern Sea Route (NSR), increasingly navigable as Arctic sea ice retreats, can substantially shorten shipping distance, but route planning is difficult because navigability and risk at a location depend on a ship's arrival time. In practice it is rarely a single-route problem: forecasts are uncertain and frequently updated, and operators must weigh many departure times and time-cost trade-offs, so a planner is queried repeatedly. Graph-search and metaheuristic methods recompute a full solution per query and become the bottleneck, while prior reinforcement learning studies on Arctic routes assume static ice and ignore arrival-time-dependent risk. This study proposes a framework in which ST-A* (Space-Time A*) generates demonstration trajectories for different time-cost weights, and DQfD (Deep Q-Learning from Demonstrations) uses them to learn weight-specific policies that produce routes for new ice conditions by inference, without the full space-time re-search. On 51 unseen test cases from 2023-2025, the learned policies reached the destination in every case with lower objective values than a baseline trained without demonstrations. Relative to ST-A*, they produced routes about 700 times faster at a mean objective 1.57 times higher, a bounded loss in per-route optimality that makes large numbers of route evaluations tractable.

키워드
Northern Sea RouteReinforcement LearningDeep Q-Learning from DemonstrationsMulti-Objective OptimizationSea Ice Risk
목차
1. 서 론
2. 문헌 연구
3. 제안 방법
    3.1 POLARIS 기반 해빙 위험도 평가
    3.2 목적함수 및 제약조건
    3.3 비용 모델
    3.4 강화학습 환경 및 리워드 설계
    3.5 Space-Time A* 기반 데모 경로 생성
    3.6 DQfD 기반 정책 학습
4. 실험 설계
    4.1 데이터 구성
    4.2 비교 대상 및 평가 지표
5. 실험 결과
    5.1 다목적 경로의 상충관계
    5.2 Space-Time A* 대비 성능
    5.3 강화학습 baseline 비교
6. 함의
    6.1 연산 효율과 경로 품질 상충관계 및 활용 방안
    6.2 데모 활용의 도달성 기여
7. 결론 및 향후 연구
Acknowledgement
References
저자
  • Seungwoo Lee(Graduate School of AI Convergence Engineering, Changwon National University) | 이승우 (국립창원대학교 대학원 인공지능융합공학과)
  • Byungheon Lee(School of Business Administration, Kwangwoon University) | 이병헌 (광운대학교 경영학부)
  • Sunghyun Sim(Department of AI Engineering, Changwon National University) | 심성현 (국립창원대학교 인공지능공학과) Corresponding author