M. Maruf

Applied Scientist at Amazon AGI · Speech & Audio Foundation Models · Multimodal AI

maruf_profile_2026.png

I’m Maruf, an Applied Scientist II at Amazon AGI Foundations, focused on building multimodal foundation models for speech and audio. My broader interests are in creating intelligent systems that can understand, generate, and interact through natural speech across languages, with an emphasis on scalable training, conversational AI, post-training, and model evaluation.

I have contributed to the pretraining, post-training, evaluation, and production delivery of Nova Sonic, Amazon’s speech-to-speech foundation model. My recent work includes designing pretraining recipes and data mixtures, building scalable data generation and evaluation pipelines, and developing post-training methods with supervised fine-tuning, Direct Preference Optimization, and reinforcement learning.

I received my Ph.D. in Computer Science from Virginia Tech, where I worked with Prof. Anuj Karpatne. My doctoral research focused on multimodal learning, vision-language models, knowledge-guided machine learning, graph representation learning, generative modeling, and AI for scientific discovery.

My research has been published at venues including NeurIPS, ICLR, CVPR, ECCV, AAAI, KDD, ICASSP, and SDM. I am broadly interested in building foundation models that can learn, reason, and interact across speech, language, vision, and other modalities.

Recent News

May 2026 Our work on Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs appeared at ICASSP 2026.
Jun 2025 Our Fish-Vista paper was published at CVPR 2025.
Apr 2025 Our work on hierarchical prototypes over the Tree of Life was published at ICLR 2025.
Dec 2024 Joined Amazon AGI Foundations as an Applied Scientist, working on speech and multimodal foundation models.
Dec 2024 Completed my Ph.D. in Computer Science at Virginia Tech.
Sep 2024 Yay! VLM4Bio is accepted in NeurIPS 2024. :tada:
Aug 2024 VLM4Bio is in arxiv!! Check it out here
Aug 2024 Our FishVista paper pre-print is available on arxiv! Check it out here.
Jul 2024 Yay :tada: PhyloDiffusion paper is accepted at ECCV 2024! Check it out: here.
Dec 2023 3 papers are accepted at the first workshop on Imageomics at AAAI, 2023. Here is my talk on “Are Pre-trained Vision Language Models (VLMs) Decent Zero-shot Predictors in Scientific Contexts?”

Selected Publications

For a complete and up-to-date list of publications, please visit my Google Scholar.

  1. Preprint
    Can We Read the Mind of Audio LLMs? A Verbalizable, Multilingual Middle-Layer Workspace
    J. Fan, J. Li, M. Maruf, and 1 more author
    2026
    Under submission
  2. ICASSP 26
    Continuous-token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
    Xinlu He, Swayambhu Nath Ray, M. Maruf, and 1 more author
    In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
  3. NeurIPS 24
    vlm4bio.jpg
    VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
    M. Maruf, A. Daw, K. S. Mehrab, and 19 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2024
  4. ICLR 25
    hcompnet.jpg
    What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
    H. B. Manogaran, M. Maruf, A. Daw, and 10 more authors
    In International Conference on Learning Representations (ICLR), 2025
  5. ECCV 24
    phyloDiffusion.jpg
    Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution
    Mridul Khurana, Arka Daw, M. Maruf, and 12 more authors
    In Proceedings of ECCV, Milan, Italy, 2024
  6. CVPR 25
    fishvista.jpg
    Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images
    K. S. Mehrab*M. Maruf*, A. Daw*, and 14 more authors
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  7. Vision
    ws3.jpg
    Beyond Discriminative Regions: Saliency Maps as Alternatives to CAMs for Weakly Supervised Semantic Segmentation
    M. Maruf, Arka Daw, Amartya Dutta, and 2 more authors
    In arXiv, 2023
  8. NeurIPS 21
    dam.jpg
    Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)
    Jie Bu*, Arka Daw*M. Maruf*, and 1 more author
    In Advances in Neural Information Processing Systems (NeurIPS), 2021
  9. KDD 21
    pidgan.jpg
    PID-GAN: A GAN Framework based on a Physics-informed Discriminator for Uncertainty Quantification with Physics
    Arka Daw*M. Maruf*, and Anuj Karpatne
    In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD), 2021
  10. SDM 20
    dns.jpg
    Maximizing cohesion and separation in graph representation learning: A distance-aware negative sampling approach
    M. Maruf, and Anuj Karpatne
    In Proceedings of the 2021 SIAM International Conference on Data Mining, 2020