Time for a serious bit of refactoring. Most of that code in those if blocks in main.py really belong in functions in some package or other. Took me a while but I realized that I might, down the road, want to, at the same time, generate different versions of a chord progression using most of the same command line arguments. Seemed to me that putting the necessary code in a function would make life a lot easier. Not to mention, main.py a lot tidier.

Not sure it is really where I need to go, but…

New Function to Generate a Complete Chord Progression

We already have a lot of functions used in this process. But, making this block of code its own function seems to make a bit of sense. Let’s see how that works out. It will be going in the em_chords package.

I plan on returning the final 1D (h-stacked) numpy array. Don’t know if it will be normalized and typed to int16. For now, not. And, I will also return the roman numeral version of the chord progression—I have been using that in the name of the wave file to which I save the musical version of the progression.

Not yet certain about what parameters should be passed to the function. But, I think I am going to pass it the verified command line argument dictionary. Though it will likely be enhanced/extended to cover all the necessary parameters for generating a chord progression. For instance I believe the time signature and the root octave should be specified in that dictionary. Not yet sure about the overtone parameters.

So that function definition will look something like the following.

def make_chord_prog(args:dict[str, Union[str, int, bool]]) -> tuple[str, npt.NDArray[np.number[int | float]]]:
  ...

Instead of assigning the dictionary values to local variables I will just use the dictionary values. Unless it makes sense to create a variable to make things work or to simplify the code. Here’s the first working version of the new function.

Added a debug print switch to the dictionary passed to the function. Had to refactor a couple other functions in the em_chords package to include and/or use a dbg parameter. Won’t bother showing those changes.

def make_chord_prog(args:dict[str, Union[str, int, bool]]) -> tuple[str, npt.NDArray[np.number[int | float]]]:
  """ Generate a numpy array of a sound track of a random or specified chord progression.
      Most necessary arguments/parameters passed in args parameter, a dictionary.

    :param args: a dictionary of all the parameters required to generate the chord progression
      keys are: whatdo,  filepath_in, filepath_wv, filepath_mid, track_key, nbr_bars, play_track,
        no_save_wv, no_save_mid, cp_type, cp_nbr, instr, r_oct, t_sig, s_rate, ot_s, ot_w, n_ot,
        dbg

    :returns: tuple consisting of the roman numeral string for the generated chord progression (str),
      numpy array containing the sampled wave data for the progression sound track as generated
      by the function
  """
  
  s_rnt, s_frm = args["track_key"]
  # should these also be passed in??
  c_dur, c_amp = 1.0, 1.0
  p_tempo = 120
  s_bt = 60 / p_tempo

  if args["dbg"]:
    print(f"\nselected key: {s_rnt} {s_frm} (octave: {args['r_oct']})")
  rn_prg, c_prg, cp_base = make_cp_chords(s_rnt, s_frm, args["r_oct"], args["t_sig"], args["cp_type"], args["cp_nbr"], dbg=args["dbg"])

  # for now use same multipliers and frequencies for overtones
  ot_mlts = emw.get_ot_mults(args["ot_s"], n_ot=args["n_ot"])
  ot_amps = emw.get_ot_amps(c_amp, args["ot_s"], h_typ=args["ot_w"], n_ot=args["n_ot"])
  if args["dbg"]:
    print(f"\tmultipliers: {ot_mlts}\n\tamplitudes: {[round(oa, 4) for oa in ot_amps]}")
  p_chds = make_chords_overtones(cp_base, ot_mlts, ot_amps,
                        c_dur=c_dur, c_amp=c_amp, s_rate=args["s_rate"],
                        do_norm=False, do_np16=False)

  # let's get a rhythm and see if we can get the sound array sorted appropriately
  cp_len = len(p_chds)
  cp_rhy = emb.make_chd_rhythm(cp_len, t_sig=args["t_sig"], n_bars=args["nbr_bars"], tempo=p_tempo, retro=False)
  cp_durs = emb.Note_durations(args["t_sig"], s_bt)
  if args["dbg"]:
    print(f"\t{cp_durs.n_dur}")

  c_amps = emw.get_amplitudes(int(args["nbr_bars"] * cp_len))
  if args["dbg"]:
    print(f"\nn_bars: {args['nbr_bars']}, cp_len: {cp_len}, len cp_rhy: {len(cp_rhy)}, len cp_durs: {len(cp_durs.n_dur)}, len c_amps: {len(c_amps)}")

  if args["dbg"]:
    print("\ncalling make_cp_sound")
  st = time.process_time()
  a_bars = make_cp_sound(p_chds, cp_rhy, cp_durs, c_amps, c_prg, nosnd=[], dbg=args["dbg"])
  et = time.process_time()
  if args["dbg"]:
    print(f"make_c_sound done: {(et-st):.4f}")

  snd = np.hstack(a_bars)
  return rn_prg, snd

And the refactored if block in main.py now looks like this. Everything from if do_sav_wav or do_w_play: down has not changed (at least yet).

  if do_mk_cprog:
    do_sav_wav = not args["no_save_wv"]

    args["r_oct"] = emw.rng.choice([1, 2, 3, 4])
    args["t_sig"] = (emw.rng.choice([3,4]), 4)
    args["s_rate"] = sample_rate
    args["ot_s"] = emw.rng.choice(["even", "odd", "seq"])
    args["ot_w"] = emw.rng.choice(["half", "saw", "sqr", "tri"])
    # should perhaps make this a random selection over some small range
    args["n_ot"] = 10
    # print debug statements in various functions
    args["dbg"] = True

    print(f"\ncommand line args are ok: {args_ok}; augmented args are:")
    for ky, val in args.items():
      print(f"{ky}: {val}")

    rn_prg, snd =  emc.make_chord_prog(args)

    if do_sav_wav or do_w_play:
      # snd = np.hstack(a_bars)
      snd = emw.normalize_wave(snd, do_typ=True)

    if do_sav_wav:
      # snd = np.hstack(p_chds)
      if args["filepath_wv"] == "":
        w_fl_nm = f"{aud_mid_dir}/{rn_prg}_{args['t_sig'][0]}-{args['t_sig'][1]}_{args['ot_s']}_{args['ot_w']}_o{args['r_oct']}_1.wav"
      else:
        w_fl_nm = args["filepath_wv"]
      # print(f"going to write to wave file: {w_fl_nm}")
      # Open a WAV file
      with wave.open(w_fl_nm, 'w') as wav_file:
        print(f"writing to wave file: {w_fl_nm}")
        # Define audio parameters
        wav_file.setnchannels(1) # Mono
        wav_file.setsampwidth(2) # Two bytes per sample
        wav_file.setframerate(sample_rate)
        # Convert the NumPy array to bytes and write it to the WAV file
        wav_file.writeframes(snd.tobytes())

    if do_w_play:
      sd.play(snd)
      sd.wait()

And a quick test.

(base) PS R:\learn\e_m_311> uv run main.py -wd m_cp -nsw -ply -nbr 6
WARNING:root:Coremltools is not installed. If you plan to use a CoreML Saved Model, reinstall basic-pitch with `pip install 'basic-pitch[coreml]'`
WARNING:root:tflite-runtime is not installed. If you plan to use a TFLite Model, reinstall basic-pitch with `pip install 'basic-pitch tflite-runtime'` or `pip install 'basic-pitch[tf]'
WARNING:root:onnxruntime is not installed. If you plan to use an ONNX Model, reinstall basic-pitch with `pip install 'basic-pitch[onnx]'`
R:\learn\e_m_311\.venv\Lib\site-packages\resampy\filters.py:50: UserWarning: pkg_resources is deprecated as an API. See https://setuptools.pypa.io/en/latest/pkg_resources.html. The pkg_resources package is slated for removal as early as 2025-11-30. Refrain from using this package or pin to Setuptools<81.
  import pkg_resources

command line args are ok: True; augmented args are:
whatdo: m_cp
filepath_in:
filepath_wv:
filepath_mid:
track_key: ('A', 'min_nat')
nbr_bars: 6
play_track: True
no_save_wv: True
no_save_mid: False
cp_type: random
cp_nbr: 3
instr: []
r_oct: 1
t_sig: (4, 4)
s_rate: 44100
ot_s: odd
ot_w: saw
n_ot: 10
dbg: True

selected key: A min_nat (octave: 1)
  scale notes: ['A1', 'B1', 'C2', 'D2', 'E2', 'F2', 'G2']
  key chords: [('A', 'minor'), ('B', 'dim'), ('C', 'major'), ('D', 'minor'), ('E', 'minor'), ('F', 'major'), ('G', 'major')]
  chord progression (roman numerals, random): I-ii-vii
  chord progression: [('A', 'minor'), ('B', 'dim'), ('G', 'major')]
    [
      A minor -> ['A1', 'C2', 'E2']
      B dim -> ['B1', 'D2', 'F2']
      G major -> ['G1', 'B1', 'D2']
    ]
        multipliers: [3, 5, 7, 9, 11, 13, 15, 17, 19, 21]
        amplitudes: [0.2359, 0.169, 0.1221, 0.099, 0.0837, 0.072, 0.0642, 0.0569, 0.0509, 0.0462]
        {'whl': 2.0, 'hlf': 1.0, 'qtr': 0.5, '8th': 0.25, '16th': 0.125, '32nd': 0.0625}

n_bars: 6, cp_len: 3, len cp_rhy: 6, len cp_durs: 6, len c_amps: 18

calling make_cp_sound
  ['qtr', 'qtr', 'hlf']
    0: qtr -> 0.5 * 0.7711 -> 0: ('A', 'minor')
    1: qtr -> 0.5 * 0.4202 -> 1: ('B', 'dim')
    2: hlf -> 1.0 * 0.3360 -> 2: ('G', 'major')
  ['qtr', 'qtr', 'hlf']
    0: qtr -> 0.5 * 0.3899 -> 0: ('A', 'minor')
    1: qtr -> 0.5 * 0.7189 -> 1: ('B', 'dim')
    2: hlf -> 1.0 * 0.2942 -> 2: ('G', 'major')
  ['qtr', 'qtr', 'hlf']
    0: qtr -> 0.5 * 0.6429 -> 0: ('A', 'minor')
    1: qtr -> 0.5 * 0.2313 -> 1: ('B', 'dim')
    2: hlf -> 1.0 * 0.6754 -> 2: ('G', 'major')
  ['qtr', 'hlf', 'qtr']
    0: qtr -> 0.5 * 0.6720 -> 0: ('A', 'minor')
    1: hlf -> 1.0 * 0.5686 -> 1: ('B', 'dim')
    2: qtr -> 0.5 * 0.5916 -> 2: ('G', 'major')
  ['qtr', 'hlf', 'qtr']
    0: qtr -> 0.5 * 0.5322 -> 0: ('A', 'minor')
    1: hlf -> 1.0 * 0.4615 -> 1: ('B', 'dim')
    2: qtr -> 0.5 * 0.7243 -> 2: ('G', 'major')
  ['hlf', 'qtr', 'qtr']
    0: hlf -> 1.0 * 0.6802 -> 0: ('A', 'minor')
    1: qtr -> 0.5 * 0.6666 -> 1: ('B', 'dim')
    2: qtr -> 0.5 * 0.6010 -> 2: ('G', 'major')
         qtr -> 0.5 * 0.6010 -> 2: ('A', 'minor')
make_c_sound done: 0.0000

I assure you, the sound track in the numpy array was played. And, a wave file was not saved to disk. Though must admit the progression really didn’t sound all that good.

Audio To/From Midi Functions

Not sure the if do_m2w: block would benefit, at least at this time, from being moved into a function. But I think I will look at doing so with some of the if do_w2m_ps: block code. Then work on the not yet coded if do_w2m_pi: block.

New Package

Going to start a new package, em_files.py, for the next couple of functions. Possibly more down the line. Have included the first function in the documentation at the top of the module.

# em_files.py: package to provide code for exploring and manipulating audio and midi files
# ver 0.1: rek, 2026.06.29, init

import time, wave

from pathlib import Path
from typing import Any, Union

from basic_pitch.inference import predict, predict_and_save, Model

""" Current Classes
"""

""" Current Functions
  convert_save_w2m(args:dict[str, Union[str, int, bool]]) -> str
"""

""" Globals
"""

do_w2m_ps

This was mainly copy and paste, followed by a lot of refactoring of variable and parameter names/values.


def convert_save_w2m(args:dict[str, Union[str, int, bool]]) -> str:
  """ Convert a wave file to midi and save to disk.
      Most necessary arguments/parameters passed in args parameter, a dictionary.

    :param args: a dictionary of all the parameters required to generate the chord progression
      keys are: whatdo,  filepath_in, filepath_wv, filepath_mid, track_key, nbr_bars, play_track,
        no_save_wv, no_save_mid, cp_type, cp_nbr, instr, r_oct, t_sig, s_rate, ot_s, ot_w, n_ot,
        fs, dbg
      requires, 'filepath_in' and 'fs'
        'filepath_in' is variable containing the path to the wave file
        'fs' is a variable pointing to initialized Fluidsynth class, used to predict midi file

    :returns: path to saved midi file
  """
  midi_fl = args["filepath_in"]
  midi_fl = midi_fl.replace(".wav", "_basic_pitch.mid")

  # process the specified file
  audio_fl = Path(args["filepath_in"])
  # save mid file to same directory wav file was in
  fl_pth = audio_fl.parent

  st = time.perf_counter()
  predict_and_save(
    [audio_fl],
    fl_pth,
    True,
    True,
    False,
    False,
    args["bpm"]
  )
  et = time.perf_counter()
  print(f"audio to midi took {(et-st):.4f} sec")

  return midi_fl

And a quick test. Refactor appropriate block in main.py.

  if do_w2m_ps:
    midi_fl = emf.convert_save_w2m(args)

And in the terminal after executing an appropriate call to main.py the following was output.

PS R:\learn\e_m_311> uv run main.py -wd w2m_ps -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav -ply
... ...
Predicting MIDI for R:\learn\e_m_311\img\I-iii-IV-I_4-4_odd_sqr_o4_1.wav...

  Creating midi...
  💅 Saved to R:\learn\e_m_311\img\I-iii-IV-I_4-4_odd_sqr_o4_1_basic_pitch.mid

  Creating midi sonification...
  🎧 Saved to R:\learn\e_m_311\img\I-iii-IV-I_4-4_odd_sqr_o4_1_basic_pitch.wav
audio to midi took 13.4046 sec
FluidSynth runtime version 2.5.4
Copyright (C) 2000-2026 Peter Hanappe and others.
Distributed under the LGPL license.
SoundFont(R) is a registered trademark of Creative Technology Ltd.

And, the midi file was played before execution ended.

do_w2m_pi

Let’s move on to the other version of converting a wave file to a midi file. This one is intended to allow for control of the instrument(s) to be used for each channel in the midi file (for now just one channel, hopefully more down the road).

In this one we use Basic Pitch’s predict method. This method returns:

  • model_output is the raw model inference output
  • midi_data is the transcribed MIDI data derived from the model_output
  • note_events is a list of note events derived from the model_output

At this point I see no use for the model_output or the note_events. But I have not done any research on either of them or their uses.

Class PrettyMIDI

Let’s have a look at some of that data. A bit of dev/test code. I do warn you, that code went through a number of iterations. Each one getting longer and generating more output.

I refer you to the code samples of pretty-midi on github for how I got to what you see below.

  if do_w2m_pi:
    if True:
      ## test code
      audio_fl = Path(args["filepath_in"]),
      m_out, m_data, n_evnts = predict(
        args["filepath_in"],
        args["bpm"]
      )
      print(f"\nmodel output: {type(m_out)}, midi data: {type(m_data)}, note events: {type(n_evnts)}")
      print(f"\nmodel output: {m_out.keys()}")
      # took me a bit to realize note events were in more or less reverse order compared
      # to notes in instrument objects
      print(f"\nnote events:")
      for e_i, e_nt in enumerate(n_evnts[-10:]):
        print(f"\t{e_i}: {e_nt[:4]}, {e_nt[4][:5]}")
      print(f"\ninstruments: {len(m_data.instruments)}, key sig chgs: {len(m_data.key_signature_changes)}, time sig chngs: {len(m_data.time_signature_changes)}, lyrics: {len(m_data.lyrics)}, text events: {len(m_data.text_events)}\n")
      for ndx, instr in enumerate(m_data.instruments):
        print(f"\t{ndx}: {instr.program}, {instr.name}, {instr.is_drum}, {len(instr.notes)}")
        for n_i, i_nt in enumerate(instr.notes[:10]):
          print(f"\t\t{n_i} -> velocity {i_nt.velocity}, pitch: {i_nt.pitch}, start: {i_nt.start}, end: {i_nt.end}")
      exit(0)

In the terminal I got the following output for one specific midi file.

PS R:\learn\e_m_311> uv run main.py -wd w2m_pi -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav

Predicting MIDI for R:\learn\e_m_311/img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav...

model output: <class 'dict'>, midi data: <class 'pretty_midi.pretty_midi.PrettyMIDI'>, note events: <class 'list'>

model output: dict_keys(['note', 'onset', 'contour'])

note events:
        0: (0.9984580498866213, 1.5209070294784581, 78, 0.8365916), [1, 1, 1, 1, 1]
        1: (0.9984580498866213, 1.497687074829932, 75, 0.7553713), [1, 1, 1, 1, 1]
        2: (0.9984580498866213, 1.497687074829932, 71, 0.6546087), [0, 1, 1, 1, 1]
        3: (0.8243083900226758, 0.9984580498866213, 73, 0.64289516), [1, 1, 1, 1, 1]
        4: (0.4992290249433107, 0.9984580498866213, 77, 0.77051383), [1, 1, 1, 1, 1]
        5: (0.4992290249433107, 0.8243083900226758, 73, 0.76075554), [1, 1, 1, 1, 1]
        6: (0.4876190476190476, 0.9868480725623583, 70, 0.6095778), [1, 1, 1, 1, 1]
        7: (0.011609977324263039, 0.4876190476190476, 73, 0.84658915), [1, 1, 1, 1, 1]
        8: (0.011609977324263039, 0.47600907029478456, 70, 0.7687936), [1, 1, 1, 1, 1]
        9: (0.011609977324263039, 0.47600907029478456, 66, 0.6229109), [1, 1, 1, 1, 1]

instruments: 1, key sig chgs: 0, time sig chngs: 0, lyrics: 0, text events: 0

        0: 4, , False, 74
                0 -> velocity 79, pitch: 66, start: 0.011609977324263039, end: 0.47600907029478456
                1 -> velocity 98, pitch: 70, start: 0.011609977324263039, end: 0.47600907029478456
                2 -> velocity 108, pitch: 73, start: 0.011609977324263039, end: 0.4876190476190476
                3 -> velocity 77, pitch: 70, start: 0.4876190476190476, end: 0.9868480725623583
                4 -> velocity 97, pitch: 73, start: 0.4992290249433107, end: 0.8243083900226758
                5 -> velocity 98, pitch: 77, start: 0.4992290249433107, end: 0.9984580498866213
                6 -> velocity 41, pitch: 58, start: 0.7314285714285714, end: 0.9287981859410431
                7 -> velocity 82, pitch: 73, start: 0.8243083900226758, end: 0.9984580498866213
                8 -> velocity 83, pitch: 71, start: 0.9984580498866213, end: 1.497687074829932
                9 -> velocity 96, pitch: 75, start: 0.9984580498866213, end: 1.497687074829932

Well enough of that.

Modify Argument Parser and Verifier

I have decided to modify the type for the command line parameter --instr from int to str. That way I can pass a prog number or a name for the new instrument. A bit more work in the code, but hopefully much more flexible. And, I updated the get_args_do_w2m() function to add that argument’s value to the args dictionary it returns.

... ...
#  in get_parser()
...
  parser.add_argument("-ift", "--instr", nargs="+", type=str, action="append",
                      help="Supply the new instrument for a given track, can be used multiple times")
... ...
# in get_args_do_w2m()
... ... 
    if a_key not in ["filepath_in", "filepath_mid", "instr", "play_track"]:
      continue
... ...
      case "instr":
        t_args[a_key] = a_val

A quick test.

PS R:\learn\e_m_311> uv run main.py -wd w2m_pi -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav -ift 0 40 -ift 1 "double bass"
... ...
command line args are ok: True
whatdo: w2m_pi
filepath_in: R:\learn\e_m_311/img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav
filepath_wv:
filepath_mid:
track_key: random
nbr_bars: 6
play_track: False
no_save_wv: False
no_save_mid: False
cp_type: random
cp_nbr: 0
instr: [['0', '40'], ['1', 'double bass']]
fs: <midi2audio.FluidSynth object at 0x000001F0699D6BD0>
bpm: <basic_pitch.inference.Model object at 0x000001F03CBD3150>

And that seems to work. Will have to deal with those numbers as strings in our code.

parse_instr_arg()

Decided on a helper function, parse_instr_arg(), currently in the em_files package. Though not sure that’s where it really belongs.

... ...
from sf_gu_gs import GS_soundfile
... ...
""" Globals
  sfgs: instantiation of soundfile class to allow prog to name for available instruments
"""

sfgs = GS_soundfile()
... ...
def parse_instr_arg(ap_instr:list[list[str, str]]) -> dict[int, tuple[int, str]] | None:
  """ Convert raw --instr arg values to suitable dictionary, keyed on channel number

    :param ap_instr: raw argparse value for argument --instr

    :returns: dictionary, keyed on channel number, with a tuple giving the new instrument
      prog number and instrument name for that channel, can be empty
  """
  c2l = {}
  for i_c, i_i in ap_instr:
    if i_c.isdecimal():
      if i_i.isdecimal():
        n_ii = int(i_i)
        c2l[int(i_c)] = (n_ii, sfgs.get_instr_nm(n_ii))
      else:
        c2l[int(i_c)] = (sfgs.get_nm2key(i_i), i_i)
  return c2l

And a quick test produced this output in the terminal.

PS R:\learn\e_m_311> uv run main.py -wd w2m_pi -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav -ift 0 40 -ift 1 "double bass"
... ...
command line args are ok: True
whatdo: w2m_pi
filepath_in: R:\learn\e_m_311/img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav
filepath_wv:
filepath_mid:
track_key: random
nbr_bars: 6
play_track: False
no_save_wv: False
no_save_mid: False
cp_type: random
cp_nbr: 0
instr: [['0', '40'], ['1', 'double bass']]
fs: <midi2audio.FluidSynth object at 0x000001CF40C2F710>
bpm: <basic_pitch.inference.Model object at 0x000001CF40F8EF50>

ch2instr: {0: (40, 'violin'), 1: (43, 'double bass')}

Back to convert_w2m_instr()

I think I can now finish coding this function and give it a quick test.

ef convert_w2m_instr(args:dict[str, Union[str, int, bool]]) -> str:
  """ Convert a wave file to midi, possibly altering instrument for specified channels.
      Save result to disk.
      Most necessary arguments/parameters passed in args parameter, a dictionary.

    :param args: a dictionary of all the parameters required to generate the chord progression
      keys are: whatdo, filepath_in, filepath_wv, filepath_mid, track_key, nbr_bars, play_track,
        no_save_wv, no_save_mid, cp_type, cp_nbr, instr, r_oct, t_sig, s_rate, ot_s, ot_w, n_ot,
        fs, dbg
      requires: 'filepath_in', 'instr', 'bpm'
        'filepath_in' is variable containing the path to the wave file
        'instr' is a list of tuples specifying channel and new instrument prog number,
          (channel, insturment prog)
        'pbm' is variable pointing to a basic pitch model

    :returns: path to saved midi file
  """

  # generate midi using a specific instrument if requested
  audio_fl = Path(args["filepath_in"])
  # save mid file to same directory wav file was in
  fl_pth = audio_fl.parent

  st = time.perf_counter()

  # get midi data
  _, m_data, _ = predict(
    audio_fl,
    args["bpm"]
  )
  # get dictionary of new instruments based on command line argument
  ch2instr = parse_instr_arg(args["instr"])
  # replace current instrument with new one, i_new
  is_chg = False
  for i_ch, instrument in enumerate(m_data.instruments):
    if i_ch in ch2instr.keys():
      i_prog = instrument.program
      i_nm = sfgs.get_instr_nm(i_prog)
      print(f"\t{instrument} -> changing ({i_prog}, {i_nm} to {ch2instr[i_ch]}")
      instrument.program = ch2instr[i_ch][0]
      is_chg = True
  # save midi file
  if is_chg:
    midi_fl = Path(f"{fl_pth}/{audio_fl.stem}_w2m_i-{ch2instr[0][1]}.mid")
  else:
    midi_fl = Path(f"{fl_pth}/{audio_fl.stem}_w2m_i-no_i_chg.mid")
  m_data.write(midi_fl)

  et = time.perf_counter()
  print(f"audio to midi took {(et-st):.4f} sec")

  return midi_fl

The if block in main.py is really pretty simple.

... ...
  if do_w2m_pi:
    midi_fl = emf.convert_w2m_instr(args)

And a quick test.

PS R:\learn\e_m_311> uv run main.py -wd w2m_pi -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav -ift 0 birds -ift 1 "double bass" -ply
... ...
Predicting MIDI for R:\learn\e_m_311\img\I-iii-IV-I_4-4_odd_sqr_o4_1.wav...
        Instrument(program=4, is_drum=False, name="") -> changing (4, tine electric piano to (123, 'birds')
audio to midi took 12.5127 sec
FluidSynth runtime version 2.5.4
Copyright (C) 2000-2026 Peter Hanappe and others.
Distributed under the LGPL license.
SoundFont(R) is a registered trademark of Creative Technology Ltd.

I thought rather than just include the wave version of the above midi, I would include the original wave file of the plain chord progression generated by the software. As well as wave version of the default midi created from that wave file. Then the version generated above with the first channel’s instrument changed.

Note, the Basic Pitch generated .wav file would not play in the browser. Guess it doesn’t like float values either. So, I did a midi to wave conversion to get a file that would play.



Play Audio/Midi Files

When creating the midi file in the code above, I have the sonify parameter to predict_and_save() set to True. As such, basic-pitch creates a wave version of the midi file it generated at the same time. But this wave file has float values. The Python wave module is not able to play such files. I didn’t want to mess around trying to code something that would convert that to 16 bit values in two channels.

But when I checked the installed packages in this virtual environment, I noticed that soundfile was installed. I used soundfile to load/read the wave file (into a numpy array), then used sounddevice to play it. Way simpler than what I had before. I had also tried using simpleaudio but couldn’t get that to work and didn’t feel like adding scipy to the environment.

  if ((not do_mk_cprog) and do_w_play):
    data, s_rate = sf.read(wv_fl)
    sd.play(data, s_rate)
    sd.wait()

And I assure you, the basic-pitch generated wave file was played by the above. And, it also plays the wave files generated by my numpy to wave code.

Compare Two Wave File Types

I was curious what the differences between the two files might be. So, I used soundfile to help out. The following code is now commented out.

    with sf.SoundFile(wv_fl) as sf_obj:
      print(f"\nSoundFile: {sf_obj}")
      print(f"\tnbr frames: {sf_obj.frames}")
      print(f"\tformat_info: {sf_obj.format_info}")
      print(f"\tsubtype_info: {sf_obj.subtype_info}")
      print(f"\tnbr sections: {sf_obj.sections}")

    exit(0)

In the terminal I got the following output when inspecting the two files.

PS R:\learn\e_m_311> uv run main.py -wd play -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav
... ...
SoundFile: SoundFile('img/I-iii-IV-I_4-4_odd_sqr_o4_1.wav', mode='r', samplerate=44100, channels=1, format='WAV', subtype='PCM_16', endian='FILE')
        nbr frames: 529200
        format_info: WAV (Microsoft)
        subtype_info: Signed 16 bit PCM
        nbr sections: 1

PS R:\learn\e_m_311> uv run main.py -wd play -fpi img/I-iii-IV-I_4-4_odd_sqr_o4_1_basic_pitch.wav
... ...
SoundFile: SoundFile('img/I-iii-IV-I_4-4_odd_sqr_o4_1_basic_pitch.wav', mode='r', samplerate=44100, channels=1, format='WAV', subtype='DOUBLE', endian='FILE')
        nbr frames: 569695
        format_info: WAV (Microsoft)
        subtype_info: 64 bit float
        nbr sections: 1

I always get a second or so of no sound when playing the basic-pitch generated files, midi or wave. Can see that that appears to be on purpose. Though certainly have no idea why. Hunt on web says, according to Google ai, that it is on purpose: “these trailing empty frames or extra length are generally caused by the default zero-padding applied to audio inputs.”

Done

I think that’s it for this post. Lots of progress (though perhaps a step or two backwards as well). Plenty of code. And a fun piece of music to finish it off. Not to mention some considerable amount of my time to get it and the related code to its current state. Hopefully this will help with our future steps.

Until we meet again, may you think things through better and sooner than I seem to.

Resources