https://www.mdu.se/

mdu.sePublications
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Reducing Manual image annotation effort using sam, dinov3 and active learning
Mälardalen University, Faculty of Engineering and Health Sciences, Department of Computer Science & Engineering. Transcom - Tele2 Kundsupport.
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

Modern computer vision systems rely heavily on annotated image datasets, but creating object-level annotations is both time-consuming and expensive. This thesis investigates how active learning can be integrated into a semi-automatic annotation workflow in order to reduce the amount of manual labeling while still achieving reliable classification performance. The proposed workflow combines the Segment Anything Model (SAM), DINOv3-based feature extraction, active learning, and a graphical interface within a unified annotation pipeline. SAM is used to extract object regions, while DINOv3 generates feature representations for classification. Active learning is then applied to iteratively select informative samples using both random sampling and uncertainty-based sampling strategies.

 

The experiments were conducted on a subset of the Microsoft COCO dataset using four experimental configurations: SAM with random sampling, SAM with uncertainty sampling, DINO with random sampling, and DINO with uncertainty sampling. Fully automatic segmentation was also compared with box-prompted segmentation to evaluate how segmentation quality affects the overall workflow. The experiments showed that classification performance improved as additional labeled samples were introduced. The guided segmentation setting achieved a mean IoU of 0.879 together with precision, recall, detection F1-score, and AP@0.5 values of 1.000. In the active learning experiments, DINO with uncertainty-based sampling achieved stronger performance during the early annotation stages, while SAM with random sampling achieved the highest final macro F1-score of 0.705 at larger annotation budgets. The experiments also showed that fully automatic segmentation generated several irrelevant object proposals, whereas box-prompted segmentation improved localization reliability and annotation quality. 

 

The findings suggest that active learning can improve annotation efficiency, but that the overall performance depends strongly on segmentation quality and feature representations.

Place, publisher, year, edition, pages
2026. , p. 42
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:mdh:diva-78612OAI: oai:DiVA.org:mdh-78612DiVA, id: diva2:2087213
Subject / course
Computer Science
Presentation
2026-06-04, Mälardalens Universitet, Universitetsplan 1, 722 20 Västerås, Västerås, 13:50 (English)
Supervisors
Examiners
Available from: 2026-08-04 Created: 2026-07-19 Last updated: 2026-08-04Bibliographically approved

Open Access in DiVA

Thesis Report(1354 kB)25 downloads
File information
File name FULLTEXT02.pdfFile size 1354 kBChecksum SHA-512
e208338488195ac796a7849abb4933babe78b9b4fd6c7ca9c97b876d3b424ae8741098ef04f546c7df89c55fef0368ccc5ffec439aa722f31ba5fd859037cd16
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Mahmoud, Fatima
By organisation
Department of Computer Science & Engineering
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 25 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 855 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf