All Stories

  1. Scalable Evaluation for Audio Identification Via Synthetic Latent Fingerprint Generation
  2. Domain-Invariant Representation Learning of Bird Sounds
  3. Computational hermeneutics: evaluating generative AI as a cultural technology
  4. RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
  5. From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
  6. Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
  7. Position Paper: Towards a Unified Representation Evaluation Framework Beyond Downstream Tasks
  8. Acoustic Prompt Tuning: Empowering Large Language Models With Audition Capabilities
  9. LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
  10. Velocity2DMs: A Contextual Modeling Approach to Dynamics Marking Prediction in Piano Performance
  11. Globally, songs and instrumental melodies are slower and higher and use more stable pitches than speech: A Registered Report
  12. A Data-Driven Analysis of Robust Automatic Piano Transcription
  13. ATGNN: Audio Tagging Graph Neural Network
  14. PiJAMA: Piano Jazz with Automatic MIDI Annotations
  15. Few-shot Class-incremental Audio Classification Using Dynamically Expanded Classifier with Self-attention Modified Prototypes
  16. Exploring Transformer’s Potential on Automatic Piano Transcription
  17. Improving Lyrics Alignment Through Joint Pitch Detection
  18. Learning Music Audio Representations Via Weak Language Supervision
  19. Automatic Quality Assessment of Digitized and Restored Sound Archives
  20. Measuring national mood with music: using machine learning to construct a measure of national valence from audio data
  21. Adaptive Scattering Transforms for Playing Technique Recognition
  22. Comparison of Feature Extraction Methods for Sound-Based Classification of Honey Bee Activity
  23. Detecting Cover Songs with Pitch Class Key-Invariant Networks
  24. Humanities and engineering perspectives on music transcription
  25. An Evaluation of Data Augmentation Methods for Sound Scene Geotagging
  26. Vocal Harmony Separation Using Time-Domain Neural Networks
  27. Violinist identification based on vibrato features
  28. MusCaps: Generating Captions for Music Audio
  29. Revisiting the Onsets and Frames Model with Additive Attention
  30. More for Less: Non-Intrusive Speech Quality Assessment with Limited Annotations
  31. Joint Multi-Pitch Detection and Score Transcription for Polyphonic Piano Music
  32. Prototypical Networks for Domain Adaptation in Acoustic Scene Classification
  33. The Effect of Spectrogram Reconstruction on Automatic Music Transcription: An Alternative Approach to Improve Transcription Accuracy
  34. Adversarial Unsupervised Domain Adaptation for Harmonic-Percussive Source Separation
  35. Development of a Speech Quality Database Under Uncontrolled Conditions
  36. Memory Controlled Sequential Self Attention for Sound Recognition
  37. Deep generative variational autoencoding for replay spoof detection in automatic speaker verification
  38. Reliable Local Explanations for Machine Listening
  39. A Study on the Transferability of Adversarial Attacks in Sound Event Classification
  40. A-CRNN: A Domain Adaptation Model for Sound Event Detection
  41. Audio Impairment Recognition using a Correlation-Based Feature Representation
  42. Modeling Plate and Spring Reverberation Using A DSP-Informed Deep Neural Network
  43. Playing Technique Recognition by Joint Time–Frequency Scattering
  44. Deep Learning for Black-Box Modeling of Audio Effects
  45. Learning and Evaluation Methodologies for Polyphonic Music Sequence Prediction With LSTMs
  46. Dataset Artefacts in Anti-Spoofing Systems: A Case Study on the ASVspoof 2017 Benchmark
  47. City Classification from Multiple Real-World Sound Scenes
  48. Investigating Kernel Shapes and Skip Connections for Deep Learning-Based Harmonic-Percussive Separation
  49. Polyphonic Sound Event and Sound Activity Detection: A Multi-Task Approach
  50. Ensemble Models for Spoofing Detection in Automatic Speaker Verification
  51. Towards Joint Sound Scene and Polyphonic Sound Event Recognition
  52. Adaptive Noise Reduction for Sound Event Detection Using Subband-Weighted NMF
  53. Optimal neural network feature selection for spatial-temporal forecasting
  54. Adapting the Quality of Experience Framework for Audio Archive Evaluation
  55. Audio-based Identification of Beehive States
  56. Automatic Transcription of Diatonic Harmonica Recordings
  57. SubSpectralNet – Using Sub-spectrogram Based Convolutional Neural Networks for Acoustic Scene Classification
  58. Automatic Music Transcription: An Overview
  59. Analysing The Predictions Of a CNN-Based Replay Spoofing Detection System
  60. ANALYSING REPLAY SPOOFING COUNTERMEASURE PERFORMANCE UNDER VARIED CONDITIONS
  61. Polyphonic Music Sequence Transduction with Meter-Constrained LSTM Networks
  62. Towards Complete Polyphonic Music Transcription: Integrating Multi-Pitch Detection and Rhythm Quantization
  63. A supervised classification approach for note tracking in polyphonic piano transcription
  64. Detection and Classification of Acoustic Scenes and Events: Outcome of the DCASE 2016 Challenge
  65. A review of manual and computational approaches for the study of world music corpora
  66. A computational study on outliers in world music
  67. Automatic Transcription of Polyphonic Vocal Music
  68. Sound event detection in synthetic audio: Analysis of the dcase 2016 task results
  69. Approaches to Complex Sound Scene Analysis
  70. Polyphonic Sound Event Tracking Using Linear Dynamical Systems
  71. On-Bird Sound Recordings: Automatic Acoustic Recognition of Activities and Contexts
  72. On the memory properties of recurrent neural models
  73. The Digital Music Lab
  74. A Morphological Model for Simulating Acoustic Scenes and Its Application to Sound Event Detection
  75. Speaker recognition with hybrid features from a deep belief network
  76. Digital music lab: A framework for analysing big music data
  77. An End-to-End Neural Network for Polyphonic Piano Music Transcription
  78. Detection of overlapping acoustic events using a temporally-constrained probabilistic model
  79. Automatic transcription of Turkish microtonal music
  80. Detection and Classification of Acoustic Scenes and Events
  81. Alternate level clustering for drum transcription
  82. A hybrid recurrent neural network for music transcription
  83. The temperament police
  84. Incremental Dataset Definition for Large Scale Musicological Research
  85. Learning motion-difference features using Gaussian restricted Boltzmann machines for efficient human action recognition
  86. Improving instrument recognition in polyphonic music through system integration
  87. Automatic transcription of pitched and unpitched sounds from polyphonic music
  88. Big Data for Musicology
  89. Detection and classification of acoustic scenes and events: An IEEE AASP challenge
  90. Automatic music transcription: challenges and future directions
  91. Multiple-instrument polyphonic music transcription using a temporally constrained shift-invariant model
  92. Temporally-Constrained Convolutive Probabilistic Latent Component Analysis for Multi-pitch Detection
  93. A temporally-constrained convolutive probabilistic model for pitch detection
  94. Joint Multi-Pitch Detection Using Harmonic Envelope Estimation for Polyphonic Music Transcription
  95. Polyphonic music transcription using note onset and offset detection
  96. Improving Music Genre Classification Using Automatically Induced Harmony Rules
  97. Auditory Spectrum-Based Pitched Instrument Onset Detection
  98. Non-Negative Tensor Factorization Applied to Music Genre Classification
  99. Computationally Efficient and Robust BIC-Based Speaker Segmentation
  100. A neural network approach to audio-assisted movie dialogue detection
  101. Systematic comparison of BIC-based speaker segmentation systems
  102. Applying Supervised Classifiers Based on Non-negative Matrix Factorization to Musical Instrument Classification
  103. Automatic Speaker Segmentation using Multiple Features and Distance Measures: A Comparison of Three Approaches