Evolving Deep Architectures: A New Blend of CNNs and Transformers Without Pre-training Dependencies
Kiiskilä, Manu; Kiiskilä, Padmasheela (2024)
Kiiskilä, Manu
Kiiskilä, Padmasheela
2024
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202409188779
https://urn.fi/URN:NBN:fi:tuni-202409188779
Kuvaus
Peer reviewed
Tiivistelmä
Modeling in computer vision is slowly moving from Convolution Neural Networks (CNNs) to Vision Transformers due to the high performance of self-attention mechanisms in capturing global dependencies within the data. Although vision transformers proved to surpass CNNs in performance and require less computational power, their need for pre-training on large-scale datasets can become burdensome. Using pre-trained models has critical limitations, including limited flexibility to adjust network structures and domain mismatches of source and target domains. To address this, a new architecture with a blend of CNNs and Transformers is proposed. SegFormer with four transformer blocks is used as an example, replacing the first two transformer blocks with two CNN modules and training from scratch. Experiments with the MS COCO dataset show a clear improvement in accuracy with C-C-T-T architecture compared to T-T-T-T architecture when trained from scratch on limited data. This project proposes an architecture modifying the SegFormer Transformer with two convolutional modules, achieving pixel accuracies of 0.6956 on MS COCO.
Kokoelmat
- TUNICRIS-julkaisut [25018]
