A Cloud-Native Multimodal Vision-Language Framework for Structured Facial Skincare Assessment and Personalized Recommendation Generation
DOI:
https://doi.org/10.65138/ijresm.v9i8.3499Abstract
Personalized skincare guidance remains inaccessible to many due to the cost and limited availability of dermatological consultations. Conventional automated skin-analysis tools rely on purpose-trained convolutional neural networks that demand large annotated datasets and continual retraining. This paper presents a Cloud-Native Multimodal Vision-Language Framework (CN-MVLF) that instead performs facial skincare assessment through prompt-engineered inference on a pre-trained multimodal large language model, avoiding bespoke model training entirely. A Groq-hosted Llama Vision model, guided by a schema-constrained prompt, returns a twenty-plus field structured JSON report covering skin-type classification, condition severity, personalized routines, and ingredient guidance. The system is deployed as a three-tier architecture, a browser client, a stateless FastAPI gateway, and a Firebase backend-as-a-service layer, with a deterministic fallback analyzer ensuring availability during API disruption. Across twenty-five trials, the framework achieves a mean response latency of 4.3 seconds, an 80% skin-type classification agreement with expert review, and a 100% structured-output parse-success rate, outperforming free-form prompting baselines while requiring no proprietary hardware or training pipeline.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 S. Manasa, Basavaraj S. Pol, K. R. Varalakshmi

This work is licensed under a Creative Commons Attribution 4.0 International License.
