Modeling conversational emotion as latent regimes using sticky factorial HDP-HMMs over multimodal valence-arousal signals. This framework recovers interpretable emotional phases, enabling context-augmented response logic during unstable affective states.
Decomposing phoneme recognition into structured articulatory features such as manner, place, and voicing to improve recognition of non-canonical speech. A cross-attention-based hierarchical multi-task architecture combines articulatory supervision with semi-supervised learning, producing more robust phoneme predictions and interpretable error patterns aligned with phonological structure.
A framework for edge-based emotion recognition using modality arbitration. Utilizing Dirichlet evidence and Dempster-Shafer theory, the system reconciles conflicting signals from speech, text, and facial micro-expressions to capture predictive uncertainty.
A graph-structured long-term memory engine designed to evolve with users over months of interaction. Mnemosyne utilizes probabilistic recall and "core summaries" to outperform both standard RAG pipelines in dialogue realism and other agentic memory systems in certain LoCoMo benchmark categories.