thumbnail

Parakrant Sarkar |





About Me


Hello, and thank you for visiting my page. I am an Assistant Professor in the Department of Computer Science and Engineering at SRM University AP, Andhra Pradesh, India. I was awarded my Ph.D. from the City University of Hong Kong on 24 December 2025 for my thesis, Learning-Based Music Enhancement for Real-World Recording Scenarios, supervised by Dr. PerMagnus Lindborg.

My research interests include neural audio effect modeling, music enhancement, hearing-aid enhancement, speech and audio processing, and machine learning. My work explores learning-based methods for improving the quality, expressiveness, and accessibility of real-world audio recordings.

Before joining SRM University AP, I worked as an Audio Researcher at Sony Computer Science Laboratories and as an Audio Research Intern at the AudioLab of the Huawei Hong Kong Research Center. I also worked as a Speech Recognition Developer at Vocera Communication.

Google scholar | DBLP

Professional Experience


● Assistant Professor | Department of Computer Science and Engineering, SRM University AP | Andhra Pradesh, India (Mar '26 - present)

● Audio Researcher | Sony Computer Science Laboratories | Tokyo, Japan (Remote) (Jul '25 - Sep '25)
    Supervisor: Dr. Taketo Akama

● Audio Research Intern | Huawei, Hong Kong SAR, China | AudioLab, Huawei Hong Kong Research Center (Apr '24 - Dec '24)
    Supervisor: Dr. Simon Liu

● Speech Recognition Developer Level 2 | Vocera, Bangalore, India | R&D Products Engineering Group (Oct '16 - Dec '20)

● Senior Scientific Officer | SRIC-IIT Kharagpur, West Bengal, India | Department of Information Technology, Govt. of India (Sep '12 - Sep '16)

Education


● City University of Hong Kong (63rd QS Rank) | Hong Kong | Ph.D. from School of Creative Media (Sep '21 - Nov '25)
    PhD Awarded: 24 December 2025
    Thesis Title: Learning-Based Music Enhancement for Real-World Recording Scenarios
    Supervisor: Dr. PerMagnus Lindborg [Mar '23 - Aug '25]
    Research Area: Neural Audio Effect Modeling, Music Enhancement, and Hearing Aid Enhancement
    Supervisor: Dr. Can Liu [Sep '21 - Mar '23]
    Research Area: Improving hands-free voice-based text editing using head-based interaction.

● Indian Institute of Technology Kharagpur | Kharagpur, India | MS from Department of Computer Science & Engineering (Jan '13 - Oct '16)
    Thesis Title: Prosody Modeling for Storytelling Style Speech Synthesis
    Supervisor: Dr. Krothapalli Sreenivasa Rao
    GPA: 9.36/10.00

● North Eastern Hill University | Meghalaya, India | Bachelor of Technology in Information Technology (Aug '08 - Aug '12)
    Major Project: Gene-gene module extraction from Gene-gene regulatory network.
    Supervisor: Dr. Swarup Roy
    GPA: 6.69/9.0, University Rank: 3rd/120

Thesis


Ph.D. thesis: Learning-Based Music Enhancement for Real-World Recording Scenarios
Ph.D. awarded on 24 December 2025. Supervisor: Dr. PerMagnus Lindborg

MS Thesis: Prosody Modeling for Storytelling Style Speech Synthesis
Abstract | Demos

Major contributions

  • Development of story TTS using a neutral TTS system in Hindi with appropriate story-specific information generation and incorporation modules. It includes design, development, and integration of story-specific prosody rule-set generation and incorporation to neutral TTS.
  • Development of story TTS using story speech corpus.
  • Modeling of story-specific pause patterns is proposed with and without discourse information.
  • Modeling of story-specific prosody (i.e., duration, intonation, and intensity) is proposed based on story genre information.
  • Publications


    ● Parakrant Sarkar, and Permagnus Lindborg, "Diff-DEQ: Differentiable Dynamic Equalization for Studio-Quality Speech Processing", in Proceedings of the IEEE 33rd European Signal Processing Conference (EUSIPCO 2025), Palermo, Itlay, 08-12 September, 2025. [Soon to be uploaded]

    ● Parakrant Sarkar, and Permagnus Lindborg, "Neural-Driven Multi-Band Proessing for Automatic Equalization and Style Transfer", in Proceedings of the IEEE 27th International COnference on Digital Audio Effects (DAFx 2025), Ancona, Itlay, 02-05 September, 2025. [Soon to be uploaded]

    ● Parakrant Sarkar, and Permagnus Lindborg, "Dynamic EQ Simplified: A Rule-Based Method for Frequency Selective Audio Processing", in 20th Workshop on Computer Music and Audio Technology (WOCMAT 2024), Ancona, Itlay, 02-05 September, 2025. [Soon to be uploaded]

    ● Kumud Tripathi, Parakrant Sarkar, and K. Sreenivasa Rao "Sentence Based Discourse Classification for Hindi Story Text-to-Speech (TTS) System", in Thirteenth International Conference on Natural Language Processing (ICON-2016), IIT (BHU), Varanasi, India, 19-20 December, 2016. pdf

    ● Parakrant Sarkar, and K. Sreenivasa Rao, "Development of Story Text-to-Speech System based on Story Genres", in Workshop on Machine Learning in Speech and Language Processing (MLSLP 2016), Google San Fransisco, USA, 13 September, 2016. pdf

    ● Parakrant Sarkar, and K. Sreenivasa Rao, "Analysis and Modeling Pauses for Synthesis of Storytelling Speech based on Discourse modes", in Proceedings of the IEEE International Conference on Contemporary Computing (IC3 2015), JIIT Noida, India, 11-13 August, 2015. pdf

    ● Parakrant Sarkar, and K. Sreenivasa Rao, "Modeling Pauses for Synthesis of Storytelling Style Speech Using Unsupervised Word Features", Second International Symposium on Computer Vision and the Internet (VisionNet'15), Procedia Computer Science, Volume 58, Pages 42-49, Kochi, India, August, 2015. pdf

    ● Parakrant Sarkar, and K. Sreenivasa Rao, "Data-Driven Pause Prediction for Synthesis of Storytelling Style Speech based on Discourse Modes", in Proceedings of the IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT 2015), IIIT Bangalore, India, 10-11 July, 2015. pdf

    ● Parakrant Sarkar, and K. Sreenivasa Rao,"Data-Driven Pause Prediction for Speech Synthesis in Storytelling Style Speech" , in Proceedings of the IEEE 21st National Conference on Communication (NCC-2015), IIT Bombay, India, 27 February to 1 March, 2015. pdf

    ● Rashmi Verma, Parakrant Sarkar, and K. Sreenivasa Rao, "Conversion of Neutral speech to Storytelling Style speech", in Proceedings of the eighth IEEE International Conference on Advances in Pattern Recognition (ICAPR 2015), ISI Kolkata, India, 04-07 January, 2015. pdf

    ● Parakrant Sarkar Arijul Haque, Arup Kumar Dutta, Gurunath Reddy M., Harikrishna D. M., Prasenjit Dhara, Rashmi Verma, N. P. Narendra, Sunil Kr. S.B., Jainath Yadav, K. Sreenivasa Rao, "Designing prosody rule-set for converting neutral TTS speech to storytelling style speech for Indian languages Bengali, Hindi and Telugu", in Proceedings of the IEEE International Conference on Contemporary Computing (IC3 2014), JIIT Noida, India, 20-24 August, 2014. pdf

    ©right 2018 Parakrant Sarkar, Last update: 14-Oct-2023