| Русский Русский | English English |
   
Главная Текущий номер
10 | 08 | 2026
10.14489/vkit.2026.07.pp.008-015

DOI: 10.14489/vkit.2026.07.pp.008-015

Иванов Н. Н., Климов М. А., Глухих И. Н., Глухих Д. И., Чернышева Т. Ю.
РАЗРАБОТКА ПРОГРАММНОГО ИНСТРУМЕНТА АВТОМАТИЗИРОВАННОЙ ГЕНЕРАЦИИ И АННОТИРОВАНИЯ ВИДЕОДАННЫХ ДЛЯ ЗАДАЧ КОМПЬЮТЕРНОГО ЗРЕНИЯ
(с. 8-15)

Аннотация. Представлен разработанный программный инструмент с открытым исходным кодом, размечающий видеоданные для обучения нейронных сетей. Научная новизна исследования заключается в создании комбинированного метода фрагментарной видеосегментации. Предлагаемый подход впервые объединяет архитектуру сегментации по промптам Segment Anything Model2 и механизм долговременной памяти XMem++ в единый автономный вычислительный конвейер. Реализованная стратегия «инициализация по требованию» позволяет динамически обновлять банк признаков объекта и эффективно устранять семантический дрейф без обращения к внешним облачным ресурсам. Приведена математическая формализация задачи трекинга и сегментации, детально описана архитектура программного комплекса. Установлено, что использование инструмента ускоряет процесс подготовки датасетов в 5–7 раз по сравнению с ручной разметкой. Рассмотрена задача, связанная с промышленной безопасностью, для наглядности применения инструмента.

Ключевые слова:  компьютерное зрение; генерация датасетов; сегментация; аннотирование данных; метод полуавтоматической аннотации; архитектура программного комплекса.


Ivanov N. N., Klimov M. A., Glukhykh I. N., Glukhykh D. I., Chernysheva T. Y.
DEVELOPMENT OF A SOFTWARE TOOL FOR AUTOMATED GENERATION AND ANNOTATION OF VIDEO DATASETS FOR COMPUTER VISION TASKS
(pp. 8-15)

Abstract. The article is devoted to solving the urgent problem of automating the markup of video data for training neural networks. The analysis of existing approaches revealed the key disadvantages of manual annotation associated with high labor costs and the risks of using foreign cloud SaaS solutions due to information security requirements and sanctions restrictions. The article presents an open source tool designed for video markup. The scientific novelty of the research lies in the creation of a combined method of fragmentary video segmentation. For the first time, the proposed approach combines the Segment Anything Model 2 segmentation architecture and the XMem++ long-term memory mechanism into a single autonomous computing pipeline. The implemented "initialization on demand" strategy allows you to dynamically update the object's feature bank and effectively eliminate semantic drift without accessing external cloud resources. The paper provides a mathematical formalization of the tracking and segmentation problem, as well as a detailed description of the architecture of the software package. An experimental evaluation conducted on a set of test video sequences confirmed the effectiveness of the proposed solution. It has been found that using the tool speeds up the dataset preparation process by 5–7 times compared to manual markup. The results obtained demonstrate the tool's readiness for use in various applications. Within the framework of the article, the task related to industrial safety is considered to illustrate the use of the tool.

Keywords: Computer vision; Dataset generation; Segmentation; Data annotation; Semi-automatic annotation method; Software architecture.

Рус

Н. Н. Иванов, М. А. Климов, И. Н. Глухих, Д. И. Глухих, Т. Ю. Чернышева (Тюменский государственный университет, Тюмень, Россия) E-mail: Этот e-mail адрес защищен от спам-ботов, для его просмотра у Вас должен быть включен Javascript  

Eng

N. N. Ivanov, M. A. Klimov, I. N. Glukhykh, D. I. Glukhykh, T. Y. Chernysheva (University of Tyumen, Tyumen, Russia) E-mail: Этот e-mail адрес защищен от спам-ботов, для его просмотра у Вас должен быть включен Javascript  

Рус

1. Discriminative Correlation Filter with Channel and Spatial Reliability / A. Lukezic, T. Vojir, L. Cehovin Zajc et al. // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), 21–26 July 2017. Honolulu, USA. P. 6309–6318.
2. Faster R-CNN detection based OpenCV CSRT tracker using drone data / X. Farhodov, O.-H. Kwon, K. W. Kang et al. // International Conference on Information Science and Communications Technologies (ICISCT 2019), 4–6 November 2019. Tashkent, Uzbekistan. P. 1–3.
3. He K., Gkioxari G., Dollár P., Girshick R. Mask R-CNN // Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017). 22–29 October 2017. Venice, Italy. P. 2961–2969.
4. Багиров М. Б., Бородина Т. Л., Карклин Т. Д., Дмитриев Д. В. Разработка системы сопровождения объектов на видеопотоке // Вестник компьютерных и информационных технологий. 2022. Т. 19. № 9(219). С. 3–13.
5. Microsoft COCO: Common Objects in Context / T.-Y. Lin, M. Maire, S. Belongie et al. // European Conference on Computer Vision (ECCV 2014). 6–12 September 2014. Zurich, SSwitzeland. Springer, Cham. P. 740–755.
6. The Pascal Visual Object Classes (VOC) Challenge / M. Everingham, L. Van Gool, C. K. I. Williams et al. // International Journal of Computer Vision. 2010. V. 88, no. 2. P. 303–338.
7. Chandan G., Jain A., Jain H. Real time object detection and tracking using deep learning and OpenCV // International Conference on Inventive Research in Computing Applications (ICIRCA 2018). 11–12 July 2018. Coimbatore, India. IEEE. P. 1305–1308.
8. Fast Segment Anything / X. Zhao, W. Ding, Y. An et al. // arXiv preprint arXiv:2306.12156. 2023. URL: https://arxiv.org/abs/2306.12156 (дата обращения: 10.12.2025).
9. Segment Anything / A. Kirillov, E. Mintun, N. Ravi et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2023). 2–6 October 2023. Paris, France. P. 4015–4026.
10. SAM 2: Segment Anything in Images and Videos / N. Ravi, V. Gabeur, Y.-T. Hu et al. // arXiv preprint arXiv:2408.00714. 2024. URL: https://arxiv.org/abs/2408.00714 (дата обращения: 10.12.2025).
11. Bekuzarov M., Bermudez A., Lee J.-Y., Li H. XMem++: Production-Level Video Segmentation from Few Annotated Frames // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2023). 2–6 October 2023. Paris, France. P. 635–644.
12. Cheng H. K., Schwing A. G. XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model // European Conference on Computer Vision (ECCV 2022). 23–27 October 2022. Tel Aviv, Israel. Cham: Springer. P. 640–658.
13. Track Anything Annotate: Video annotation and dataset generation of computer vision models / N. Ivanov, M. Klimov, D. Glukhikh et al. // arXiv preprint arXiv:2505.17884. 2025. URL: https://arxiv.org/abs/2505.17884 (дата обращения: 10.12.2025).
14. Redmon J., Divvala S., Girshick R., Farhadi A. You Only Look Once: Unified, Real-Time Object Detection // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016). 27–30 June 2016. Las Vegas, USA. P. 779–788.

Eng

1. Lukezic, A., Vojir, T., Cehovin Zajc, L., et al. (2017). Discriminative correlation filter with channel and spatial reliability. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017) (pp. 6309–6318). Honolulu, HI, United States.
2. Farhodov, X., Kwon, O.-H., Kang, K. W., et al. (2019). Faster R CNN detection based OpenCV CSRT tracker using drone data. In International Conference on Information Science and Communications Technologies (ICISCT 2019) (pp. 1–3). Tashkent, Uzbekistan.
3. He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017) (pp. 2961–2969). Venice, Italy.
4. Bagirov, M. B., Borodina, T. L., Karklin, T. D., & Dmitriev, D. V. (2022). Development of a system for tracking objects in a video stream. Vestnik komp'yuternykh i informatsionnykh tekhnologiy, 19(9), 3–13. [in Russian language].
5. Lin, T.-Y., Maire, M., Belongie, S., et al. (2014). Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV 2014) (pp. 740–755). Springer, Cham.
6. Everingham, M., Van Gool, L., Williams, C. K. I., et al. (2010). The Pascal Visual Object Classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303–338.
7. Chandan, G., Jain, A., & Jain, H. (2018). Real time object detection and tracking using deep learning and OpenCV. In International Conference on Inventive Research in Computing Applications (ICIRCA 2018) (pp. 1305–1308). Coimbatore, India.
8. Zhao, X., Ding, W., An, Y., et al. (2023). Fast Segment Anything. arXiv preprint. https://arxiv.org/abs/2306.12156
9. Kirillov, A., Mintun, E., Ravi, N., et al. (2023). Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2023) (pp. 4015–4026). Paris, France.
10. Ravi, N., Gabeur, V., Hu, Y.-T., et al. (2024). SAM 2: Segment Anything in images and videos. arXiv preprint. https://arxiv.org/abs/2408.00714
11. Bekuzarov, M., Bermudez, A., Lee, J.-Y., & Li, H. (2023). XMem++: Production level video segmentation from few annotated frames. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2023) (pp. 635–644). Paris, France.
12. Cheng, H. K., & Schwing, A. G. (2022). XMem: Long term video object segmentation with an Atkinson Shiffrin memory model. In European Conference on Computer Vision (ECCV 2022) (pp. 640–658). Springer, Cham.
13. Ivanov, N., Klimov, M., Glukhikh, D., et al. (2025). Track Anything Annotate: Video annotation and dataset generation of computer vision models. arXiv preprint. https://arxiv.org/abs/2505.17884
14. Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016) (pp. 779–788). Las Vegas, NV, United States.

Рус

Статью можно приобрести в электронном виде (PDF формат).

Стоимость статьи 700 руб. (в том числе НДС 20%). После оформления заказа, в течение нескольких дней, на указанный вами e-mail придут счет и квитанция для оплаты в банке.

После поступления денег на счет издательства, вам будет выслан электронный вариант статьи.

Для заказа скопируйте doi статьи:

10.14489/vkit.2026.07.pp.008-015

и заполните  форму 

Отправляя форму вы даете согласие на обработку персональных данных.

.

 

Eng

This article  is available in electronic format (PDF).

The cost of a single article is 700 rubles. (including VAT 20%). After you place an order within a few days, you will receive following documents to your specified e-mail: account on payment and receipt to pay in the bank.

After depositing your payment on our bank account we send you file of the article by e-mail.

To order articles please copy the article doi:

10.14489/vkit.2026.07.pp.008-015

and fill out the  form  

 

.

 

 

 
Поиск
Баннер
Rambler's Top100 Яндекс цитирования