Sign language translator app
USD 250–750
About the project
We are building a bidirectional, real-time web application that translates between spoken English and Nigerian Sign Language (NSL) using an animated 3D avatar. The platform will support multiple use cases including live conversations, website/video accessibility, and extended hardware integration. Core Features 1. Bidirectional Real-Time Translation • Signer Mode: User signs in NSL to the camera → AI recognizes signs → outputs natural English text + speech. • Speaker Mode: User speaks English into the microphone → AI translates to NSL → 3D animated character signs in real time. 2. Additional Key Features • Captions Tool: Real-time captions for spoken content and sign language interpretation. • Live Translation on Websites & Embedded Videos: Embeddable widget for live translation on any website or video player (captions + avatar signing). • Embedded API: Public API for developers to integrate the translation service into third-party apps and platforms. • Smart Glasses Support: Optimized output mode for smart glasses (low-latency pose data stream or simplified avatar view). • Custom Avatar Upload: Users can upload their own avatar/character, which the system will rig and use as the signing avatar. 3. Animation & Output Requirements • Real-time generative skeleton-driven animation (Transformer/LSTM-based or equivalent) targeting sub-second latency and 30 FPS fluid motion. • Support for custom uploaded avatars: automatic rigging, bone mapping, and pose application (Quaternions/Euler angles). • Smoothing, co-articulation, facial expressions, and non-manual markers essential for natural NSL. 4. Data & Model Strategy • No pre-trained NSL model exists yet. • Plan to collect ~1,000 videos of diverse NSL signers for training both recognition and generation models. • Design the system with transfer learning/fine-tuning in mind for the custom dataset. 5. Technical Priorities • Fully web-based, low end-to-end latency (<1–2 seconds). • MediaPipe (or equivalent) for real-time pose estimation. • Privacy-focused, on-device inference where possible. • Responsive UI with camera/mic access and clean embedding options. • Extensible architecture for future languages and features.
Skills required
This job is listed on Freelancer.com. AiZity aggregates listings for discovery only and is not the employer. To bid or apply, use the button in the sidebar.