SignHunter: ASL Video Lookup Tool
- SignHunter is a video-based, example-driven ASL lookup tool that retrieves and ranks sign matches from a curated sign bank for human confirmation.
- It enables streamlined linguistic annotation by integrating with SignStream, displaying the top five candidate matches for efficient sign identification.
- The system employs an AI recognition backend achieving around 81% top-1 and over 92% top-5 accuracy, enhancing practical ASL dictionary access.
"SignHunter" (Editor's term) denotes the video-based lookup capability for American Sign Language (ASL) described in "New Capability to Look Up an ASL Sign from a Video Example" (Neidle et al., 2024). It is designed for the case in which an unknown ASL sign cannot be found reliably through conventional dictionary organization by English glosses or through manual specification of articulatory properties. The system accepts a video example of a sign—either a webcam recording of a citation-form sign or a clip segmented from continuous signing—returns the five most likely sign matches in decreasing order of likelihood, and links a confirmed result to the corresponding entry in the ASLLRP Sign Bank. The same lookup capability is also integrated into SignStream(R) to support linguistic annotation of ASL video data (Neidle et al., 2024).
1. Problem formulation and lexicographic motivation
The system addresses a basic but long-standing problem in ASL lookup: most ASL dictionaries are organized based on English glosses, even though there is no convention for assigning English-based glosses to ASL signs and no 1-1 correspondence between ASL signs and English words. The problem is especially acute when the target sign’s meaning is unknown or when its possible English translation(s) are not known. Dictionaries that support search by articulatory properties such as handshapes, locations, and movement properties provide an alternative, but the paper characterizes this process as cumbersome and not always successful (Neidle et al., 2024).
Within that framing, the system is not presented as a general-purpose automatic translation system. Its stated purpose is sign lookup by example. The design choice to return a ranked set of candidates rather than a single label follows directly from that purpose: the user is expected to confirm the intended sign after reviewing likely matches. This suggests a retrieval-oriented conception of recognition in which ranking quality is operationally more important than a single forced prediction.
2. Lookup workflow and interaction model
The end-to-end workflow begins with submission of a short video of a single sign through a web interface. After processing, the uploaded video is deleted immediately for privacy, while the system retains aggregate statistics on whether the correct answer appeared as the 1st, 2nd, 3rd, 4th, or 5th choice. The recognition backend then analyzes the video and returns the five most likely sign matches, sorted by likelihood. The user can play the source video alongside the candidate matches, inspect variants when a sign has multiple lexical or phonological variants, and confirm the intended sign. If none of the five is correct, the user can return to the ASLLRP Sign Bank and search by other means. Once a match is confirmed, the user is taken directly to that sign’s Sign Bank entry (Neidle et al., 2024).
The interaction model is therefore explicitly hybrid. Recognition proposes candidates; human inspection resolves ambiguity. The paper’s examples emphasize this ranked-review behavior, including cases in which the correct sign is the leftmost or near-leftmost candidate and cases involving variant disambiguation. A plausible implication is that usability depends less on perfect top-1 classification than on the frequency with which the correct sign appears in a short candidate list.
3. ASLLRP Sign Bank as lexical substrate
The ASLLRP Sign Bank is the lexical resource underlying the lookup system. It is not treated as a mere inventory of labels. Rather, it stores distinct sign productions, including lexical variants, together with video examples and metadata. Each distinct sign production receives a unique English-based gloss label, but the authors state explicitly that these glosses are approximations and are not intended to function as authoritative translations. The Sign Bank supports search by gloss text, by related English words, by handshape information, and now by video example (Neidle et al., 2024).
For each sign entry, users can view sign occurrences and play either the “sign video” itself or a broader “sign clip” containing surrounding context from continuous signing. The paper’s COMPARE example shows an entry linked to multiple occurrences and related English words such as “compare,” “comparison,” and “contrast.” The AMBULANCE example illustrates the handling of variants by presenting additional options for confirmation.
The currently recognized inventory comprises roughly 2,360 distinct signs and similar variants. It includes lexical signs, loan signs, numbers, and compounds. It does not currently include fingerspelled forms, classifiers, index signs, or gestures as standalone searchable entries, except when such forms occur as parts of compounds. This scope is substantial for practical lookup, but the paper states that coverage remains far from exhaustive.
4. Recognition backend and reported performance
The paper does not provide a full derivation of the machine-learning model in the overview article, but it attributes the recognition backend to the AI approach of Zhou et al. (2024). The model was trained on about 98,000 consistently labeled video examples compiled from ASLLRP resources. The reported evaluation emphasizes both top-1 and top-5 performance, consistent with the system’s use as a human-confirmed lookup tool rather than a fully automatic annotation mechanism (Neidle et al., 2024).
| Setting | Top-1 | Top-5 |
|---|---|---|
| Citation-form signs | 81.21% | 95.36% |
| Signs segmented from continuous signing | 80.39% | 92.96% |
Earlier in the introduction, nearly identical rounded values are reported: 80.8% top-1 and 95.2% top-5 for citation-form signs, and 80.4% top-1 and 93.0% top-5 for segmented signs. In either reporting style, the central empirical point is the same: the correct answer appears in the top five very often, which is what the authors identify as making the lookup workflow practical.
The paper further notes that recognition may be lower for ASL learner productions, because such productions can differ from the proficient signer data used in training. It also situates the system relative to prior video-based lookup work such as Gloss-Finder. A cited user study found that 10 hearing learners preferred video-based lookup over list-based or parameter-based lookup for unknown signs, but Gloss-Finder’s target sign appeared among the top 12 candidates only 66% of the time. The comparison is used to argue that substantially higher top-5 retrieval makes the present system more suitable for practical lookup.
5. Integration into SignStream(R) annotation workflows
A major aspect of the system is its integration into SignStream(R), where the lookup function becomes part of a linguistic annotation workflow rather than a standalone dictionary interface. In SignStream versions before 3.5.0, users could already search the Sign Bank by gloss and handshape and insert resulting sign data into an annotation. The new version adds the ability to search by a video clip directly from the annotation interface. The annotator sets the boundaries of an unknown sign in a video, launches the “DAI Video Search” module, reviews the top five matches, and, after confirmation, automatically inserts the sign’s gloss and associated morphophonological information into the annotation (Neidle et al., 2024).
The paper describes this insertion step as including the sign’s gloss, handshape, sign type, and other features in utterance-level annotation with minimal manual typing. The intended effect is increased efficiency and consistency in ASL annotation, especially when annotators encounter signs whose gloss is unknown or uncertain. This suggests that the lookup capability is not only a lexicographic aid but also an instrument for normalizing annotation practice across datasets and annotators.
6. Limitations, scope conditions, and relation to adjacent research
The paper identifies several limitations. Inventory coverage remains restricted to about 2,360 signs and similar variants, leaving many ASL forms outside the searchable space. Fingerspelled signs, classifiers, index signs, and gestures are not yet searchable as standalone entries. The system is presently tied to the ASLLRP Sign Bank after confirmation, although the authors note that it could in principle be linked to other dictionaries or sign resources. The authors also state that they plan user studies with ASL learners in collaboration with Matt Huenerfauth’s group at RIT and intend to continue collecting usage statistics to monitor real-world lookup success rates (Neidle et al., 2024).
Two recurrent misconceptions are addressed by the system’s design. First, the English-based glosses in the Sign Bank are not claimed to be authoritative translations; they are approximating labels for distinct sign productions. Second, the system is not positioned as fully automatic annotation. Its ranked top-5 output presupposes human confirmation.
The topic is also distinct from adjacent lines of research that involve signs but solve different technical problems. RF sensor-based ASL trigger-sign recognition for user interfaces addresses wake-sign detection in mixed activity/signing streams, uses a 77 GHz FMCW radar, and reports a trigger sign detection rate of 98.9% together with 92% sequential recognition accuracy; that problem concerns trigger detection under privacy and low-light constraints rather than lexical lookup from submitted sign video (Kurtoglu et al., 2021). Likewise, 3D air-signature work that uses pen tip and pen tail trajectories with SliTCNN concerns air-signature recognition and forgery robustness rather than ASL dictionary access (Atreya et al., 2024). These contrasts clarify the niche occupied by SignHunter: example-based retrieval of ASL lexical items, grounded in a curated sign bank and coupled to annotation infrastructure.