GenASL

GenASL

open_source

GenASL by AWS translates speech or text into human-like ASL avatar videos using Amazon Bedrock, Transcribe, and SageMaker. An open-source solution for accessibility.

About

GenASL is an open-source AWS reference architecture and solution that uses cutting-edge generative AI to automatically translate speech and text into expressive, human-like American Sign Language (ASL) avatar videos. Developed and published by AWS, it leverages a powerful stack of managed cloud services to deliver end-to-end sign language accessibility at scale. The solution accepts audio, video, or plain text as input. Audio and video are first processed by Amazon Transcribe to extract text, which is then converted into ASL gloss (a notational system for sign language) using large foundation models hosted on Amazon Bedrock. The gloss output is mapped to precomputed ASL pose animations — derived from the ASL Lexicon Video Dataset containing over 3,300 signs — and rendered into a final avatar video using RTMPose-based pose estimation. The workflow is orchestrated via AWS Step Functions and exposed through Amazon API Gateway, with a mobile-friendly web front end distributed via AWS Amplify. Storage is handled by Amazon S3, with Amazon DynamoDB used for the NoSQL data layer and Amazon Cognito managing secure user access. GenASL is ideal for organizations building accessible digital experiences, media companies adding ASL overlays to content, educational platforms serving the deaf community, and developers looking to integrate sign language translation into their own applications. Because it is built on scalable, serverless AWS infrastructure, it can handle everything from small use cases to enterprise-scale deployments.

Key Features

  • Speech & Text to ASL Translation: Converts audio, video, or plain text input into ASL gloss notation using Amazon Transcribe and large foundation models on Amazon Bedrock.
  • Human-like Avatar Video Generation: Generates realistic ASL avatar animations using a dataset of over 8,000 poses derived from the ASL Lexicon Video Dataset and RTMPose-based estimation.
  • Serverless AWS Architecture: Orchestrates the full pipeline with AWS Step Functions, Lambda, API Gateway, and S3 for a scalable, low-maintenance deployment.
  • Mobile-Friendly Web App: Delivers a responsive web front end via AWS Amplify, allowing users to upload input and receive ASL videos directly on mobile or desktop devices.
  • Multi-format Input Support: Accepts audio recordings, video files, or raw text, making it flexible for a wide range of content and communication workflows.

Use Cases

  • Adding ASL captions or avatar overlays to online video content to improve accessibility for deaf viewers.
  • Translating educational lectures, course materials, or training videos into ASL for inclusive learning platforms.
  • Enabling customer service portals and help centers to deliver support content in ASL.
  • Assisting healthcare providers in communicating with deaf patients through AI-generated ASL explanations.
  • Building accessibility tools for broadcasters, news organizations, or government agencies that must provide sign language interpretation.

Pros

  • Accessibility-Focused Design: Purpose-built to bridge communication gaps for the deaf and hard-of-hearing community, enabling truly inclusive digital experiences.
  • Built on Robust AWS Infrastructure: Leverages enterprise-grade, managed AWS services for reliability, security, and scalability without requiring custom infrastructure management.
  • Open-Source & Extensible: Published as an open AWS solution, allowing developers to inspect, customize, and extend the architecture to fit specific use cases.
  • Multiple Input Formats: Supports audio, video, and text inputs, making integration straightforward for diverse content pipelines and applications.

Cons

  • Requires AWS Expertise: Deploying and operating the solution requires familiarity with multiple AWS services, which may present a barrier for non-technical teams.
  • AWS Service Costs Apply: While the solution itself is open source, running it incurs usage costs for services like Amazon Bedrock, Transcribe, SageMaker, and others.
  • Limited to American Sign Language: Currently designed specifically for ASL and does not support other sign languages such as BSL, Auslan, or ISL out of the box.

Frequently Asked Questions

What is GenASL?

GenASL is an AWS-published, open-source generative AI solution that automatically translates speech or text into American Sign Language (ASL) avatar videos, making content accessible to the deaf and hard-of-hearing community.

What AWS services does GenASL use?

GenASL uses Amazon Bedrock (foundation models), Amazon Transcribe (speech-to-text), Amazon SageMaker, AWS Step Functions, AWS Lambda, Amazon API Gateway, Amazon S3, Amazon DynamoDB, Amazon Cognito, Amazon EC2, and AWS Amplify.

What input formats does GenASL support?

GenASL accepts audio files, video files, and plain text as input. Audio and video are automatically transcribed before being processed into ASL animations.

Is GenASL free to use?

The GenASL solution architecture is open source and free to deploy. However, running it requires an AWS account, and you will incur costs based on your usage of the underlying AWS services.

Who is GenASL designed for?

GenASL is designed for developers and organizations that want to add ASL accessibility to their applications, including media companies, educational platforms, healthcare providers, and any business serving the deaf or hard-of-hearing community.

Reviews

No reviews yet. Be the first to review this tool.

Alternatives

See all