WEKO3
アイテム
An Implementation of Parallel 1-D FFT on the K computer
https://nied-repo.bosai.go.jp/records/6384
https://nied-repo.bosai.go.jp/records/6384740d2193-4411-4299-b656-c6883a1ac723
| Item type | researchmap(1) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 公開日 | 2023-09-20 | |||||||||||
| タイトル | ||||||||||||
| 言語 | en | |||||||||||
| タイトル | An Implementation of Parallel 1-D FFT on the K computer | |||||||||||
| 言語 | ||||||||||||
| 言語 | eng | |||||||||||
| 著者 |
Daisuke Takahashi
× Daisuke Takahashi
× Atsuya Uno
× Mitsuo Yokokawa
|
|||||||||||
| 抄録 | ||||||||||||
| 内容記述タイプ | Other | |||||||||||
| 内容記述 | In this paper, we propose an implementation of a parallel one-dimensional fast Fourier transform (FFT) on the K computer. The proposed algorithm is based on the six-step FFT algorithm, which can be altered into the recursive six-step FFT algorithm to reduce the number of cache misses. The recursive six-step FFT algorithm improves performance by utilizing the cache memory effectively. We use the recursive six-step FFT algorithm to implement the parallel one-dimensional FFT algorithm. The performance results of one-dimensional FFTs on the K computer are reported. We successfully achieved a performance of over 18 TFlops on 8192 nodes of the K computer (82944 nodes, 128 GFlops/node, 10.6 PFlops peak performance) for a 2(41)-point FFT. | |||||||||||
| 言語 | en | |||||||||||
| 書誌情報 |
en : 2012 IEEE 14TH INTERNATIONAL CONFERENCE ON HIGH PERFORMANCE COMPUTING AND COMMUNICATIONS & 2012 IEEE 9TH INTERNATIONAL CONFERENCE ON EMBEDDED SOFTWARE AND SYSTEMS (HPCC-ICESS) p. 344-350, 発行日 2012 |
|||||||||||
| 出版者 | ||||||||||||
| 言語 | en | |||||||||||
| 出版者 | IEEE COMPUTER SOC | |||||||||||
| ISSN | ||||||||||||
| 収録物識別子タイプ | EISSN | |||||||||||
| 収録物識別子 | 2576-3512 | |||||||||||
| DOI | ||||||||||||
| 関連識別子 | 10.1109/HPCC.2012.53 | |||||||||||