{"id":6942,"date":"2021-07-16T16:30:00","date_gmt":"2021-07-16T16:30:00","guid":{"rendered":"http:\/\/desres20.netornot.at\/?p=6942"},"modified":"2021-07-16T17:47:44","modified_gmt":"2021-07-16T17:47:44","slug":"ml-sample-generator-project-phase-2-pt2","status":"publish","type":"post","link":"http:\/\/desres20.netornot.at\/?p=6942","title":{"rendered":"ML Sample Generator Project | Phase 2 pt2"},"content":{"rendered":"\n<style>\n.justify{text-align: justify}\n<\/style>\n\n\n\n<h2>Autoencoder Results<\/h2>\n\n\n\n<p class=\"justify has-medium-font-size\">As mentioned in the post before I have trained nine autoencoders to (re)produce snare drum samples. For easier comparison I have visualized the results below. Each image shows the location of all ~7500 input samples.<\/p>\n\n\n\n<h5>Rectified Linear Unit<\/h5>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr.png\" alt=\"\" class=\"wp-image-6943\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Small relu ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-1.png\" alt=\"\" class=\"wp-image-6944\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-1.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-1-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr-1-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Medium relu ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full is-style-default\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebr.png\" alt=\"\" class=\"wp-image-6945\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebr.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebr-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebr-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Big relu ae<\/figcaption><\/figure>\n\n\n\n<p class=\"justify has-medium-font-size\">All three graphics portray how the samples are mostly close together but some are very far out. A continuous representation is with all three models not possible. Reducing the latent vector\u2019s maximum on both axes definitely helps, but even then the resulting samples are not too pleasing to hear. The small network has clicks in the beginning and generates very silent but noisy tails after the initial impact. The medium network includes some quite okay samples but moving around in the latent space often &nbsp; produces &nbsp; similar&nbsp; but&nbsp; less &nbsp; pronounced issues as the small network. And the big network produces the best sounding samples but has no continuous changes.<\/p>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aesr2021-7-16-193030.wav\"><\/audio><figcaption>Clicky small relu sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemr2021-7-16-193147.wav\"><\/audio><figcaption>Noisy medium relu sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebr2021-7-16-193243.wav\"><\/audio><figcaption>Quite good big relu sample<\/figcaption><\/figure>\n\n\n\n<h5>Hyperbolic Tangent<\/h5>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aest.png\" alt=\"\" class=\"wp-image-6951\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aest.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aest-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aest-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Small tanh ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemt.png\" alt=\"\" class=\"wp-image-6952\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemt.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemt-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemt-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Medium tanh ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebt.png\" alt=\"\" class=\"wp-image-6953\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebt.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebt-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebt-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Big tanh ae<\/figcaption><\/figure>\n\n\n\n<p class=\"justify has-medium-font-size\">These three networks each produce different patterns with a cluster at (0|0). The similarities between the medium and the big network lead me to believe that there is a smooth transition between random noise, to forming small clusters, to turning 45\u00b0 clockwise and refining the clusters when increasing the number of trainable parameters. Just like the relu version, the reproduced audio samples of the small network contain clicks. The samples are however much better. The medium sized network is the best one out of all the trained models. It produces&nbsp; mostly&nbsp; good&nbsp; samples&nbsp; and has a continuous latent space. One issue is however that there are still some clicky areas in the latent space. The big network is the second best overall as it mostly lacks a continuous latent space as well. The produced audio samples are however very pleasing to hear and resemble the originals quite well.<\/p>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aest2021-7-16-193716.wav\"><\/audio><figcaption>Clicky small tanh sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aemt2021-7-16-193821.wav\"><\/audio><figcaption>Close-to-original medium tanh sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebt2021-7-16-193941.wav\"><\/audio><figcaption>Close-to-original big tanh sample<\/figcaption><\/figure>\n\n\n\n<h5>Sigmoid<\/h5>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aess.png\" alt=\"\" class=\"wp-image-6957\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aess.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aess-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aess-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Small sig ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aems.png\" alt=\"\" class=\"wp-image-6958\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aems.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aems-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aems-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Medium sig ae<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" loading=\"lazy\" width=\"640\" height=\"480\" src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebs.png\" alt=\"\" class=\"wp-image-6959\" srcset=\"http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebs.png 640w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebs-300x225.png 300w, http:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebs-360x270.png 360w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><figcaption>Big sig ae<\/figcaption><\/figure>\n\n\n\n<p class=\"justify has-medium-font-size\">This group shows a clear tendency to cluster up the more trainable parameters exist. While in the above two groups the medium and the big network produced better results, in this case the small network is by far the best. The big network delivers primarily noisy audio samples and the medium network very noisy ones as well but they are better identifiable as snare drum sounds. The small network has by far the closest sounds to the originals but produces clicks at the beginning as well.<\/p>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aess2021-7-16-194231.wav\"><\/audio><figcaption>Clicky small sigmoid sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aems2021-7-16-194313.wav\"><\/audio><figcaption>Noisy medium sigmoid sample<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/desres20.netornot.at\/wp-content\/uploads\/2021\/07\/aebs2021-7-16-19446.wav\"><\/audio><figcaption>Super noisy big sigmoid sample<\/figcaption><\/figure>\n\n\n\n<p class=\"justify has-medium-font-size\">In the third part of this series we will take a closer look at the other models.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Autoencoder Results As mentioned in the post before I have trained nine autoencoders to (re)produce snare drum samples. For easier comparison I have visualized the results below. Each image shows the location of all ~7500 input samples. Rectified Linear Unit All three graphics portray how the samples are mostly close together but some are very<\/p>\n<footer class=\"entry-footer index-entry\">\n<div class=\"post-social pull-left\"><a href=\"https:\/\/www.facebook.com\/sharer\/sharer.php?u=http%3A%2F%2Fdesres20.netornot.at%2F%3Fp%3D6942\" target=\"_blank\" class=\"social-icons\"><i class=\"fa fa-facebook\" aria-hidden=\"true\"><\/i><\/a><a href=\"https:\/\/twitter.com\/home?status=http%3A%2F%2Fdesres20.netornot.at%2F%3Fp%3D6942\" target=\"_blank\" class=\"social-icons\"><i class=\"fa fa-twitter\" aria-hidden=\"true\"><\/i><\/a><a href=\"https:\/\/www.linkedin.com\/shareArticle?mini=true&#038;url=http%3A%2F%2Fdesres20.netornot.at%2F%3Fp%3D6942&#038;title=ML+Sample+Generator+Project+%7C+Phase+2+pt2\" target=\"_blank\" class=\"social-icons\"><i class=\"fa fa-linkedin\" aria-hidden=\"true\"><\/i><\/a><\/div>\n<p class=\"link-more\"><a href=\"http:\/\/desres20.netornot.at\/?p=6942\" class=\"more-link\">Continue reading <span class=\"meta-nav\">\u2192<\/span><\/a><\/p>\n<\/footer>\n","protected":false},"author":37,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[4],"tags":[246,240,212],"_links":{"self":[{"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/posts\/6942"}],"collection":[{"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/users\/37"}],"replies":[{"embeddable":true,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=6942"}],"version-history":[{"count":4,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/posts\/6942\/revisions"}],"predecessor-version":[{"id":6966,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=\/wp\/v2\/posts\/6942\/revisions\/6966"}],"wp:attachment":[{"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=6942"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=6942"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/desres20.netornot.at\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=6942"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}