I’m Maruf, an Applied Scientist II at Amazon AGI Foundations, focused on building multimodal foundation models for speech and audio. My broader interests are in creating intelligent systems that can understand, generate, and interact through natural speech across languages, with an emphasis on scalable training, conversational AI, post-training, and model evaluation.
I have contributed to the pretraining, post-training, evaluation, and production delivery of Nova Sonic, Amazon’s speech-to-speech foundation model. My recent work includes designing pretraining recipes and data mixtures, building scalable data generation and evaluation pipelines, and developing post-training methods with supervised fine-tuning, Direct Preference Optimization, and reinforcement learning.
I received my Ph.D. in Computer Science from Virginia Tech, where I worked with Prof. Anuj Karpatne. My doctoral research focused on multimodal learning, vision-language models, knowledge-guided machine learning, graph representation learning, generative modeling, and AI for scientific discovery.
My research has been published at venues including NeurIPS, ICLR, CVPR, ECCV, AAAI, KDD, ICASSP, and SDM. I am broadly interested in building foundation models that can learn, reason, and interact across speech, language, vision, and other modalities.
Recent News
| May 2026 | Our work on Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs appeared at ICASSP 2026. |
|---|---|
| Jun 2025 | Our Fish-Vista paper was published at CVPR 2025. |
| Apr 2025 | Our work on hierarchical prototypes over the Tree of Life was published at ICLR 2025. |
| Dec 2024 | Joined Amazon AGI Foundations as an Applied Scientist, working on speech and multimodal foundation models. |
| Dec 2024 | Completed my Ph.D. in Computer Science at Virginia Tech. |
| Sep 2024 | Yay! VLM4Bio is accepted in NeurIPS 2024. |
| Aug 2024 | VLM4Bio is in arxiv!! Check it out here |
| Aug 2024 | Our FishVista paper pre-print is available on arxiv! Check it out here. |
| Jul 2024 | Yay |
| Dec 2023 | 3 papers are accepted at the first workshop on Imageomics at AAAI, 2023. Here is my talk on “Are Pre-trained Vision Language Models (VLMs) Decent Zero-shot Predictors in Scientific Contexts?” |
Selected Publications
For a complete and up-to-date list of publications, please visit my Google Scholar.
- PreprintCan We Read the Mind of Audio LLMs? A Verbalizable, Multilingual Middle-Layer Workspace2026Under submission
- ICASSP 26Continuous-token Diffusion for Speaker-Referenced TTS in Multimodal LLMsIn IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026