DOI QR코드

DOI QR Code

B2B Recommendation Model Using Multimodal Learning for Startup-Buyer Matching

이미지-텍스트 멀티모달을 활용한 국내 스타트업–해외 바이어간 B2B 추천 모델

  • 오소진 (한남대학교 경영정보학과) ;
  • 김재경 (한남대학교 경영정보학과)
  • Received : 2025.03.13
  • Accepted : 2025.04.14
  • Published : 2025.04.30

Abstract

This study proposes a B2B recommendation model utilizing a multimodal learning approach to enhance product matching between domestic startups and international buyers. Traditional recommendation systems primarily rely on textual data, often leading to information loss, especially in visually significant industries. To address this limitation, this study integrates image and text modalities using the CLIP model, enabling more accurate and context-aware recommendations. A dataset was constructed by extracting product images and descriptions from the BuyKorea platform, focusing on five major product categories. The model was trained and evaluated using Precision@K, MAP@K and NDCG@K, demonstrating a top-5 precision of 82.05%. The results confirm that multimodal learning effectively enhances recommendation quality compared to text-only approaches. This research contributes to advancing recommendation models for global B2B platforms by leveraging image-text representations.

본 연구는 국내 스타트업과 해외 바이어 간의 제품 매칭을 향상시키기 위해 멀티모달 학습 접근법을 활용한 B2B 추천 모델을 제안한다. 기존 추천 시스템은 주로 텍스트 데이터에 의존하기 때문에, 특히 시각적 요소가 중요한 산업에서는 정보 손실이 발생할 가능성이 크다. 이러한 한계를 해결하기 위해 본 연구에서는 CLIP 모델을 활용하여 이미지와 텍스트 모달리티를 통합함으로써 보다 정확하고 문맥을 반영한 추천을 가능하게 했다. 연구를 위해 BuyKorea 플랫폼에서 제품 이미지와 설명을 추출하여 5대 주요 제품 카테고리에 대한 데이터셋을 구축했다. 모델의 성능은 Precision@K, MAP@K, NDCG@K 지표를 사용하여 평가했으며, Top-5 Precision이 82.05%를 기록했다. 결과적으로 멀티모달 학습이 텍스트 기반 접근법보다 추천 품질을 효과적으로 향상시킴을 확인할 수 있었다. 본 연구는 이미지-텍스트 표현을 활용하여 글로벌 B2B 플랫폼의 추천 모델 발전에 기여할 것으로 기대된다.

Keywords

Acknowledgement

This work was supported by 2022 Hannam University Research Fund.

References

  1. Baltrušaitis, T., Ahuja, C. and Morency, L. P. (2018). Multimodal Machine Learning: A Survey and Taxonomy, IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423-443. https://doi.org/10.1109/TPAMI.2018.2798607
  2. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I. and Amodei, D. (2020). Language Models are Few-Shot Learners, Advances in Neural Information Processing Systems, https://dl.acm.org/doi/abs/10.5555/3495724.3495883
  3. Chia, P. J., Attanasio, G., Bianchi, F., Terragni, S., Magalhaes, A. R., Goncalves, D. and Tagliabue, J. (2022). Contrastive Language and Vision Learning of General Fashion Concepts, Scientific Reports, 12(1), https://www.nature.com/articles/s41598-022-23052-9.
  4. Conde, M. V. and Turgutlu, K. (2021). Clip-art: Contrastive Pre-training for Fine-grained Art Classification, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR) Workshops, Jun. 19-25, Nashville, TN, USA, pp. 3956-3960. https://doi.org/10.1109/CVPRW53098.2021.00444
  5. Devlin, J., Chang, M. W., Lee, K. and Toutanova, K. (2019). Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 2-7, Minneapolis, MN, USA, 1, pp. 4171-4186. https://aclanthology.org/N19-1423/
  6. Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S. and Chen, W. (2022). Lora: Low-rank Adaptation of Large Language Models, ICLR, 1(2), 3. https://arxiv.org/pdf/2106.09685v1/1000 106.09685v1/1000
  7. Kang, W. C., Cheng, D. Z., Yao, T., Yi, X., Chen, T., Hong, L. and Chi, E. H. (2021). Learning to Embed Categorical Features without Embedding Tables for Recommendation, Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, Aug. 14-18, Virtual Event, Singapore, pp. 840-850. https://dl.acm.org/doi/abs/10.1145/3447548.3467304
  8. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R. and Amodei, D. (2020). Scaling Laws for Neural Language Models, ArXiv, https://arXiv:2001.08361.
  9. Kim, J. (2023). A Study on Fine-Tuning and Transfer Learning to Construct Binary Sentiment Classification Model in Korean Text, Journal of Korea Society of Industrial Information Systems, 28(5), 15-30. https://doi.org/10.11627/jksie.2023.46.4.015
  10. Lee, K., Kim, M., Hong S. G. and Roh, S. (2024). Development of a Ranking System for Tourist Destination Using BERT-based Semantic Search, Journal of Korea Society of Industrial Information Systems, 29(4), 91-103. https://doi.org/10.9723/JKSIIS.2024.29.4.091
  11. Lee, S. J. and Lee, H. C. (2007). A Study on the Relation of Top-N Recommendation and the Rank Fitting of Prediction Value through a Improved Collaborative Filtering Algorithm, Journal of Korea Society of Industrial Information Systems, 12(4), 65-72.
  12. Oh, J. (2023). The Technological Evolution and Research Trends of Generative AI: Focusing on Language Models, KISDI AI Outlook, 37-52. https://kiss.kstudy.com/Detail/Ar?key=4045236
  13. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G. and Sutskever, I. (2021). Learning Transferable Visual Models from Natural Language Supervision, Proceedings of the 38th International Conference on Machine Learning (ICML), Jul. 18-24, Virtual Event, pp. 8748-8763. https://doi.org/10.48550/arXiv.2103.00020
  14. Sun, R., Cao, X., Zhao, Y., Wan, J., Zhou, K., Zhang, F. and Zheng, K. (2020). Multi-modal Knowledge Graphs for Recommender Systems, Proceedings of the 29th ACM International Conference on Information & Knowledge Management, Oct. 19-23, Virtual Event, Ireland, pp. 1405-1414. https://dacm.org/doi/abs/10.1145/3340531.3411947
  15. Sun, F., Liu, J., Wu, J., Pei, C., Lin, X., Ou, W. and Jiang, P. (2019). BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer, Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Nov. 3-7, Beijing, China, pp. 1441-1450.
  16. Sun, Y., Xu, Q., Wang, Z. and Huang, Q. (2023). When Measures are Unreliable: Imperceptible Adversarial Perturbations Toward Top-K Multi-Label Learning, Proceedings of the 31st ACM International Conference on Multimedia, Oct. 29-Nov. 3, Ottawa, ON, Canada, pp. 1515-1526.
  17. Wu, S., Sun, F., Zhang, W., Xie, X. and Cui, B. (2022). Graph Neural Networks in Recommender Systems: A Survey, ACM Computing Surveys, 55(5), 1-37. https://doi.org/10.1145/3535101
  18. Wu, X., Xia, Y., Zhu, J., Wu, L., Xie, S. and Qin, T. (2022). A Study of BERT for Context-aware Neural Machine Translation, Machine Learning, 111(3), 917-935. https://doi.org/10.1007/s10994-021-06070-y
  19. Yoon, Y. L., Yoon, Y., Nam, H. and Choi, J. (2021). Buyer-Supplier Matching in Online B2B Marketplace: An Empirical Study Of Small- and Medium-Sized Enterprises (SMEs), Industrial Marketing Management, 90-100. https://doi.org/10.1016/j.indmarman.2020.12.010
  20. Zhang, Y., Ai, Q., Chen, X. and Croft, W. B. (2017). Joint Representation Learning for Top-n Recommendation with Heterogeneous Information Sources, Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, Nov. 6-10, Singapore, Singapore, pp. 1449-1458. https://dl.acm.org/doi/10.1145/3132847.3132892
  21. Zhang, Y., Kang, B., Hooi, B., Yan, S. and Feng, J. (2023). Deep Long-tailed Learning: A Survey, IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 10795-10816. https://doi.org/10.1109/TPAMI.2023.3268118