Real-time AI-powered audio-reactive image transformation
Capture real-time audio from microphone using Web Audio API. The system continuously analyzes the incoming sound stream.
Extract audio features: volume, pitch, energy, and spectral centroid. Each characteristic maps to different visual effects.
User uploads an image that serves as the base for transformation. GPT-4o vision analyzes the image content and style.
GPT-4o generates contextual prompts based on image analysis and current audio features, maintaining the essence of the original.
DALL-E 3 creates new images based on smart prompts, ensuring audio-reactive transformations relate to the original upload.
Continuous transformations every 8 seconds in live mode, with additional real-time post-processing effects using Sharp.