Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Built with AssemblyAI - Real-time Speech-to-Image Generation

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
481
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

In the AssemblyAI project, students at ASU HACKML 2022 utilized the Core Transcription API to create real-time speech-to-image generation. They reproduced elements of DALL-E 2's zero-shot capabilities with a simpler model. The project integrates Machine Learning models and web interface framework with AssemblyAI API, enabling corrective language modeling. The inspiration came from Open AI's paper on Zero-Shot Text-to-Image Generation. The build consists of real-time audio transcription using the AssemblyAI API, HTML/CSS for client-side interface, Node.js and Express for server hosting, pretrained models running in parallel with client and server, and Selenium to pass messages between components. The main takeaways include impressive improvements in audio transcription tools, the potential of less data for similar results as larger models, and the impact of changing pretraining paradigms. Future directions involve incorporating knowledge graphs for semantic correctness checks, associating natural language with image sequences, and generating videos from resultant image vectors.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 4 1,174 339 115 -7%
LLM 1 37 11 9 -53%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.