Publication:
Vision Transformer-based driver fatigue detection in real-time embedded systems

dc.contributor.coauthorEyidoğan, F.
dc.contributor.coauthorKaya, Z. S.
dc.contributor.coauthorBuğuş, A.
dc.contributor.coauthorElmi, S.
dc.contributor.departmentGraduate School of Sciences and Engineering
dc.contributor.kuauthorElmi, Soheila
dc.contributor.schoolcollegeinstituteGRADUATE SCHOOL OF SCIENCES AND ENGINEERING
dc.date.accessioned2026-07-17T08:30:28Z
dc.date.issued2026
dc.description.abstractDriver fatigue is a major contributor to traffic accidents, yet reliable early warning remains challenging under real-world and embedded constraints. This study presents a camera-only driver-fatigue detection pipeline that couples lightweight face tracking (MediaPipe) with a single Vision Transformer (ViT) classifier ( $$\approx$$ 86M parameters) to infer alert vs. drowsy states from facial cues. The model is trained on a multi-source dataset of $$\sim$$ 4,000 images aggregated from seven public fatigue datasets and evaluated on a held-out test split, achieving 95.0% accuracy, 93.8% F1-score, and ROC-AUC of 0.986. On Raspberry Pi 5, end-to-end inference runs at 55.6 ms/frame (18.0 FPS). Compared with a detection-driven YOLO baseline using the same labeling protocol and test split, the proposed ViT improves accuracy (95.0% vs. 93.2%) and ROC-AUC (0.986 vs. 0.979) while trading throughput (18.0 vs. 28.0 FPS), clarifying the accuracy–latency balance for deployment. Real-world validation with 12 participants (22–47 years) across day and night sessions ( $$\sim$$ 9 hours total) shows high agreement with manual annotations (92.8%), supporting operational feasibility. By combining a single-backbone ViT inference pipeline with explicit embedded benchmarking, matched baseline comparison, and in-vehicle validation, the proposed system provides a deployment-oriented evaluation package for camera-only fatigue monitoring under embedded ITS constraints.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuN/A
dc.description.versionPublished Version
dc.identifier.ScopusPercentile63
dc.identifier.ScopusQuartileQ2
dc.identifier.WoSPercentile29.4
dc.identifier.WoSQuartileQ3
dc.identifier.doi10.1007/s13177-026-00673-2
dc.identifier.eissn1868-8659
dc.identifier.embargoN/A
dc.identifier.issn1348-8503
dc.identifier.scopus2-s2.0-105040530233
dc.identifier.urihttp://doi.org/10.1007/s13177-026-00673-2
dc.identifier.urihttps://hdl.handle.net/20.500.14288/33519
dc.identifier.wos001778501600001
dc.keywordsDriver fatigue
dc.keywordsImage processing
dc.keywordsVision transformer
dc.keywordsReal-time systems
dc.keywordsTraffic safety
dc.keywordsDeep learning
dc.keywordsClassifier (UML)
dc.keywordsInference
dc.keywordsPipeline (software)
dc.keywordsProtocol (science)
dc.keywordsWarning system
dc.keywordsThroughput
dc.keywordsTransformer
dc.languageeng
dc.publisherSpringer
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofInternational Journal of Intelligent Transportation Systems Research
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectTransportation science
dc.subjectTechnology
dc.titleVision Transformer-based driver fatigue detection in real-time embedded systems
dc.typeJournal Article
dspace.entity.typePublication
relation.isOrgUnitOfPublication3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isOrgUnitOfPublication.latestForDiscovery3fc31c89-e803-4eb1-af6b-6258bc42c3d8
relation.isParentOrgUnitOfPublication434c9663-2b11-4e66-9399-c863e2ebae43
relation.isParentOrgUnitOfPublication.latestForDiscovery434c9663-2b11-4e66-9399-c863e2ebae43

Files