A Cloud-Native Multimodal Vision-Language Framework for Structured Facial Skincare Assessment and Personalized Recommendation Generation

Authors

  • S. Manasa Department of Computer Science and Engineering, R. L. Jalappa Institute of Technology, Kodigehalli, India
  • Basavaraj S. Pol Department of Computer Science and Engineering, R. L. Jalappa Institute of Technology, Kodigehalli, India
  • K. R. Varalakshmi Department of Computer Science and Engineering, R. L. Jalappa Institute of Technology, Kodigehalli, India

DOI:

https://doi.org/10.65138/ijresm.v9i8.3499

Abstract

Personalized skincare guidance remains inaccessible to many due to the cost and limited availability of dermatological consultations. Conventional automated skin-analysis tools rely on purpose-trained convolutional neural networks that demand large annotated datasets and continual retraining. This paper presents a Cloud-Native Multimodal Vision-Language Framework (CN-MVLF) that instead performs facial skincare assessment through prompt-engineered inference on a pre-trained multimodal large language model, avoiding bespoke model training entirely. A Groq-hosted Llama Vision model, guided by a schema-constrained prompt, returns a twenty-plus field structured JSON report covering skin-type classification, condition severity, personalized routines, and ingredient guidance. The system is deployed as a three-tier architecture, a browser client, a stateless FastAPI gateway, and a Firebase backend-as-a-service layer, with a deterministic fallback analyzer ensuring availability during API disruption. Across twenty-five trials, the framework achieves a mean response latency of 4.3 seconds, an 80% skin-type classification agreement with expert review, and a 100% structured-output parse-success rate, outperforming free-form prompting baselines while requiring no proprietary hardware or training pipeline.

127 84

Downloads

Download data is not yet available.

Downloads

Published

07-08-2026

Issue

Section

Articles

How to Cite

[1]
S. Manasa, B. S. Pol, and K. R. Varalakshmi, “A Cloud-Native Multimodal Vision-Language Framework for Structured Facial Skincare Assessment and Personalized Recommendation Generation”, IJRESM, vol. 9, no. 8, pp. 4–10, Aug. 2026, doi: 10.65138/ijresm.v9i8.3499.

Most read articles by the same author(s)