Acoustic scene classification - DCASE2016:
'via Blog this'
Friday, January 15, 2016
Thursday, January 14, 2016
Voice Isolation: When Noise Reduction Meets Deep Learning | EE Times
http://www.eetimes.com/author.asp?section_id=36&doc_id=1328589
Voice Isolation: When Noise Reduction Meets Deep Learning | EE Times: "Voice Isolation: When Noise Reduction Meets Deep Learning"
A new approach to old problems -- with the help of deep neural networks -- may make background noise a thing of the past.
Voice Isolation: When Noise Reduction Meets Deep Learning | EE Times: "Voice Isolation: When Noise Reduction Meets Deep Learning"
A new approach to old problems -- with the help of deep neural networks -- may make background noise a thing of the past.
Sometimes it's easy to forget that a smartphone is also a telephone. With all the fantastic functions and features, somehow people have grown accustomed to the occasional dropped syllable and garbled sounds that make us repeat ourselves time and again.A recent article in Scientific American suggests that the fault is with the service providers. It's true that bandwidth is definitely a factor, but even when there is a relatively good connection, throw in a noisy environment like a coffee shop or morning traffic and communication starts to break down. A new approach to old problems -- with the help of deep neural networks -- may make background noise a thing of the past.
The voice band -- good enough?
Since the invention of the telegraph almost two centuries ago, there has been nearly exponential improvement in bandwidth, mobility, speed, and reliability. That said, a key aspect of voice telecommunications has lagged behind: the quality and intelligibility of transmitted voice. Very early on, the standard for human voice transmission was set as the "voice band" located between 300 Hz and 3.3 kHz (to put this in perspective, the natural frequency span of human voice during speech ranges from about 50 Hz to nearly 10 kHz).
Since the invention of the telegraph almost two centuries ago, there has been nearly exponential improvement in bandwidth, mobility, speed, and reliability. That said, a key aspect of voice telecommunications has lagged behind: the quality and intelligibility of transmitted voice. Very early on, the standard for human voice transmission was set as the "voice band" located between 300 Hz and 3.3 kHz (to put this in perspective, the natural frequency span of human voice during speech ranges from about 50 Hz to nearly 10 kHz).
Apparently, this was satisfactory for landline usage in quiet settings, and generations of phone users came to expect poor call quality. When these standards were carried over for cellphone audio quality, and with the added woes of spotty network coverage and connection dropouts, cellphone users' expectations for call quality fell even lower.
Extending the frequency (for better or for worse)
Now that there are about about as many cellphone subscriptions as there are people on earth, one would think that there really shouldn't be any more technological excuses for poor voice quality. New standards branded as HD Voice and VoLTE promise the eventual extension of voice transmission frequency range up to 7 kHz. An IEEE Spectrum article from September 2014 gave an instructive, in-depth analysis of the causes of lousy voice quality, and placed hope in the deployment of these new technologies. Their implementation requires new hardware and new networks, which will be overcome in time, but broadening the voice band does nothing to solve the other major challenge preventing great sounding calls -- in fact, HD Voice and its relatives may actually make the problem worse!
Now that there are about about as many cellphone subscriptions as there are people on earth, one would think that there really shouldn't be any more technological excuses for poor voice quality. New standards branded as HD Voice and VoLTE promise the eventual extension of voice transmission frequency range up to 7 kHz. An IEEE Spectrum article from September 2014 gave an instructive, in-depth analysis of the causes of lousy voice quality, and placed hope in the deployment of these new technologies. Their implementation requires new hardware and new networks, which will be overcome in time, but broadening the voice band does nothing to solve the other major challenge preventing great sounding calls -- in fact, HD Voice and its relatives may actually make the problem worse!
Noise and the boundless crusade for its cancellation
Nearly half of all phone users today employ their mobile phones as their primary voice connection (a number sure to grow). Mobile phones, by design, are used in many different environments: in planes, trains, and automobiles; at sporting events, offices, factories, and shopping centers; on playgrounds and (yeah, thatguy) in public restrooms.
Nearly half of all phone users today employ their mobile phones as their primary voice connection (a number sure to grow). Mobile phones, by design, are used in many different environments: in planes, trains, and automobiles; at sporting events, offices, factories, and shopping centers; on playgrounds and (yeah, thatguy) in public restrooms.
Just think of the noises you might encounter walking through an urban downtown, near a construction site, or in an airport lounge. While the narrow range of the current voice band standard impairs the quality of the voice that is transmitted, it also automatically filters out any noise that may be present in higher frequency bands. By doubling the frequency span of the voice band, HD Voice and relatives increase the environmental noise power that is transmitted and, ironically, can make voice quality and intelligibility worse in everyday use cases.
The noise challenges facing cellphone users are a far cry from the relatively stable noise environments that exist around landlines and, by-and-large, noise reduction technologies have not caught up. From commonly-used phase cancellation, to techniques using multiple microphones or statistical properties and mathematical assumptions about environmental noise, each attempt to isolate and cancel out noise has its deficiencies. Either some of the noise gets through or the voice suffers from audio artifacts.
Voice isolation instead of noise cancellation
The engineering team at Cypher took a different tack when developing its noise reduction technology. Instead of formulating the problem as one of capturing a signal and then eliminating the noise, they considered its mathematical dual: how to characterize and isolate the speech components of a noisy signal. Rather than trying to defend against all of the possible noise types -- an impossible task -- Cypher concentrates on elucidating and extracting common elements of human speech.
The engineering team at Cypher took a different tack when developing its noise reduction technology. Instead of formulating the problem as one of capturing a signal and then eliminating the noise, they considered its mathematical dual: how to characterize and isolate the speech components of a noisy signal. Rather than trying to defend against all of the possible noise types -- an impossible task -- Cypher concentrates on elucidating and extracting common elements of human speech.
At the core of this approach is a sophisticated deep learning methodology that identifies mathematical descriptors, which can be used in training neural networks for audio pattern recognition. The deep learning stage takes place offline using a large database of human speech. The goal of the learning is to identify and separate human speech from any environmental noise.
The result is a deep neural network that can identify in real-time precisely when and where in an audio signal the human voice is present. Despite its broad and robust pattern recognition capabilities, this deep neural network is fast enough and compact enough to run in software on the CEVA-TeakLite-4 DSP. The neural network also guides other algorithmic components of Cypher's patented technology as they isolate the person speaking from all other sources of noise -- even other nearby human speakers. Once the desired voice has been extracted, post-processing modules enhance the voice signal and remove artifacts created in the background noise elimination process. The final output has a balanced, full sound as close to the original speaker as possible. To experience the clarity, visit Cypher's website for demonstrations of this cutting edge technology in a variety of different environments.
Eran Belaish serves as CEVA's Marketing Manager of Audio and Voice Product Line, overseeing Audio and Voice processing, Android interfaces, wearable devices, and Wireless Audio. Prior to this position, Eran served as CEVA's Senior Compiler Group Leader responsible for managing all compiler-related research and development, and before that he held several engineering and management positions at CEVA since 2003. Eran holds a B.Sc. in Electrical Engineering and Computer Science from Tel-Aviv University.
Dr. Erik Sherwood serves as Cypher's Chief Scientist. Erik is an applied mathematician with expertise in dynamical systems, computational neuroscience, scientific computing, algorithm design, statistics, and machine learning. Prior to joining Cypher, he worked in academia and held teaching, research, and faculty positions in the mathematics departments of Cornell University, Boston University, and, most recently, the University of Utah. Erik studied at Princeton, the University of Bremen (Germany), Cambridge University, and Cornell University. He earned an AB with honors in mathematics and certificates in applied mathematics, computer science, and German from Princeton University, and MS and PhD degrees in applied mathematics from Cornell University.
'via Blog this'
Saturday, January 9, 2016
Machine learning / data science 面经以及一些总结
关键字: data science,data scientist,machine learning,面经
发信站: BBS 未名空间站 (Fri Jan 8 11:52:21 2016, 美东)
Source: http://www.mitbbs.com/article_t/JobHunting/33120253.html
本着国人互助以及传递正能量的真理,发一下我个人找工作过程中整理的machine
learning相关面经以及一些心得总结。楼主的背景是fresh CS PhD in computer
vision and machine learning, 非牛校。
已经有前辈总结过很多machine learning的面试题(传送门: http://www.mitbbs.com/article/JobHunting/32808273_0.html),此帖是对其的补充,有一小部分是重复的。面经分两大块:machine learning questions 和 coding questions.
Machine learning related questions:
- Discuss how to predict the price of a hotel given data from previous
years
- SVM formulation
- Logistic regression
- Regularization
- Cost function of neural network
- What is the difference between a generative and discriminative algorithm
- Relationship between kernel trick and dimension augmentation
- What is PCA projection and why it can be solved by SVD
- Bag of Words (BoW) feature
- Nonlinear dimension reduction (Isomap, LLE)
- Supervised methods for dimension reduction
- What is naive Bayes
- Stochastic gradient / gradient descent
- How to predict the age of a person given everyone’s phone call history
- Variance and Bias (a very popular question, watch Andrew’s class)
- Practices: When to collect more data / use more features / etc. (watch
Andrew’s class)
- How to extract features of shoes
- During linear regression, when using each attribute (dimension)
independently to predict the target value, you get a positive weight for
each attribute. However, when you combine all attributes to predict, you get
some large negative weights, why? How to solve it?
- Cross Validation
- Reservoir sampling
- Explain the difference among decision tree, bagging and random forest
- What is collaborative filtering
- How to compute the average of a data stream (very easy, different from
moving average)
- Given a coin, how to pick 1 person from 3 persons with equal probability.
Coding related questions:
- Leetcode: Number of Islands
- Given the start time and end time of each meeting, compute the smallest
number of rooms to host these meetings. In other words, try to stuff as many
meetings in the same room as possible
- Given an array of integers, compute the first two maximum products(乘积)
of any 3 elements (O(nlogn))
- LeetCode: Reverse words in a sentence (follow up: do it in-place)
- LeetCode: Word Pattern
- Evaluate a formula represented as a string, e.g., “3 + (2 * (4 - 1) )”
- Flip a binary tree
- What is the underlying data structure for JAVA hashmap? Answer: BST, so
that the keys are sorted.
- Find the lowest common parent in a binary tree
- Given a huge file, each line of which is a person’s name. Sort the names
using a single computer with small memory but large disk space
- Design a data structure to quickly compute the row sum and column sum of
a sparse matrix
- Design a wrapper class for a pointer to make sure this pointer will
always be deleted even if an exception occurs in the middle
- My Google onsite questions: http://www.mitbbs.com/article_t/JobHunting/33106617.html
面试的一点点心得:
最重要的一点,我觉得是心态。当你找了几个月还没有offer,并且看到别人一直在版
上报offer的时候,肯定很焦虑甚至绝望。我自己也是,那些报offer的帖子,对我来说
都是负能量,绝对不去点开看。这时候,告诉自己四个字:继续坚持。我相信机会总会
眷顾那些努力坚持的人,付出总有回报。
machine learning的职位还是很多的,数学好的国人们优势明显,大可一试, 看到一些
帖子说这些职位主要招PhD,这个结论可能有一定正确性。但是凭借我所遇到的大部分
面试题来看,个人认为MS或者PhD都可以。MS的话最好有一些学校里做project的经验。
仔细学习Andrew Ng在Coursera上的 machine learning课,里面涵盖很多面试中的概念
和题目。虽然讲得比较浅显,但对面试帮助很大。可以把video的速度调成1.5倍,节省
时间。
如果对一些概念或算法不清楚或者想加深理解,找其他的各种课件和视频学习,例如
coursera,wiki,牛校的machine learning课件。
找工作之前做好对自己的定位。要弄清楚自己想做什么,擅长做什么,如何让自己有竞
争力,然后取长补短(而不是扬长避短)。
感觉data scientist对coding的要求没有software engineer那么变态。不过即便如此
,对coding的复习也不应该松懈。
我个人觉得面试machine learning相关职位前需要熟悉的四大块:
Classification:
Logistic regression
Neural Net (classification/regression)
SVM
Decision tree
Random forest
Bayesian network
Nearest neighbor classification
Regression:
Neural Net regression
Linear regression
Ridge regression (add a regularizer)
Lasso regression
Support Vector Regression
Random forest regression
Partial Least Squares
Clustering:
K-means
EM
Mean-shift
Spectral clustering
Hierarchical clustering
Dimension Reduction:
PCA
ICA
CCA
LDA
Isomap
LLE
Neural Network hidden layer
最后祝各位好运。那些还在继续找工作的亲们,坚持住,加油!
发信站: BBS 未名空间站 (Fri Jan 8 11:52:21 2016, 美东)
Source: http://www.mitbbs.com/article_t/JobHunting/33120253.html
本着国人互助以及传递正能量的真理,发一下我个人找工作过程中整理的machine
learning相关面经以及一些心得总结。楼主的背景是fresh CS PhD in computer
vision and machine learning, 非牛校。
已经有前辈总结过很多machine learning的面试题(传送门: http://www.mitbbs.com/article/JobHunting/32808273_0.html),此帖是对其的补充,有一小部分是重复的。面经分两大块:machine learning questions 和 coding questions.
Machine learning related questions:
- Discuss how to predict the price of a hotel given data from previous
years
- SVM formulation
- Logistic regression
- Regularization
- Cost function of neural network
- What is the difference between a generative and discriminative algorithm
- Relationship between kernel trick and dimension augmentation
- What is PCA projection and why it can be solved by SVD
- Bag of Words (BoW) feature
- Nonlinear dimension reduction (Isomap, LLE)
- Supervised methods for dimension reduction
- What is naive Bayes
- Stochastic gradient / gradient descent
- How to predict the age of a person given everyone’s phone call history
- Variance and Bias (a very popular question, watch Andrew’s class)
- Practices: When to collect more data / use more features / etc. (watch
Andrew’s class)
- How to extract features of shoes
- During linear regression, when using each attribute (dimension)
independently to predict the target value, you get a positive weight for
each attribute. However, when you combine all attributes to predict, you get
some large negative weights, why? How to solve it?
- Cross Validation
- Reservoir sampling
- Explain the difference among decision tree, bagging and random forest
- What is collaborative filtering
- How to compute the average of a data stream (very easy, different from
moving average)
- Given a coin, how to pick 1 person from 3 persons with equal probability.
Coding related questions:
- Leetcode: Number of Islands
- Given the start time and end time of each meeting, compute the smallest
number of rooms to host these meetings. In other words, try to stuff as many
meetings in the same room as possible
- Given an array of integers, compute the first two maximum products(乘积)
of any 3 elements (O(nlogn))
- LeetCode: Reverse words in a sentence (follow up: do it in-place)
- LeetCode: Word Pattern
- Evaluate a formula represented as a string, e.g., “3 + (2 * (4 - 1) )”
- Flip a binary tree
- What is the underlying data structure for JAVA hashmap? Answer: BST, so
that the keys are sorted.
- Find the lowest common parent in a binary tree
- Given a huge file, each line of which is a person’s name. Sort the names
using a single computer with small memory but large disk space
- Design a data structure to quickly compute the row sum and column sum of
a sparse matrix
- Design a wrapper class for a pointer to make sure this pointer will
always be deleted even if an exception occurs in the middle
- My Google onsite questions: http://www.mitbbs.com/article_t/JobHunting/33106617.html
面试的一点点心得:
最重要的一点,我觉得是心态。当你找了几个月还没有offer,并且看到别人一直在版
上报offer的时候,肯定很焦虑甚至绝望。我自己也是,那些报offer的帖子,对我来说
都是负能量,绝对不去点开看。这时候,告诉自己四个字:继续坚持。我相信机会总会
眷顾那些努力坚持的人,付出总有回报。
machine learning的职位还是很多的,数学好的国人们优势明显,大可一试, 看到一些
帖子说这些职位主要招PhD,这个结论可能有一定正确性。但是凭借我所遇到的大部分
面试题来看,个人认为MS或者PhD都可以。MS的话最好有一些学校里做project的经验。
仔细学习Andrew Ng在Coursera上的 machine learning课,里面涵盖很多面试中的概念
和题目。虽然讲得比较浅显,但对面试帮助很大。可以把video的速度调成1.5倍,节省
时间。
如果对一些概念或算法不清楚或者想加深理解,找其他的各种课件和视频学习,例如
coursera,wiki,牛校的machine learning课件。
找工作之前做好对自己的定位。要弄清楚自己想做什么,擅长做什么,如何让自己有竞
争力,然后取长补短(而不是扬长避短)。
感觉data scientist对coding的要求没有software engineer那么变态。不过即便如此
,对coding的复习也不应该松懈。
我个人觉得面试machine learning相关职位前需要熟悉的四大块:
Classification:
Logistic regression
Neural Net (classification/regression)
SVM
Decision tree
Random forest
Bayesian network
Nearest neighbor classification
Regression:
Neural Net regression
Linear regression
Ridge regression (add a regularizer)
Lasso regression
Support Vector Regression
Random forest regression
Partial Least Squares
Clustering:
K-means
EM
Mean-shift
Spectral clustering
Hierarchical clustering
Dimension Reduction:
PCA
ICA
CCA
LDA
Isomap
LLE
Neural Network hidden layer
最后祝各位好运。那些还在继续找工作的亲们,坚持住,加油!
Friday, January 8, 2016
Apple Buys AI Startup That Reads Emotions in Faces | WIRED
Apple Buys AI Startup That Reads Emotions in Faces | WIRED:
Reference:
http://www.bloomberg.com/news/articles/2016-01-07/apple-buys-startup-that-sees-what-s-behind-your-smile
'via Blog this'
In December, when WIRED spoke to Andrew Moore, the dean of computer science at Carnegie Mellon, he said that 2016 would be the year that machines learn to grasp human emotions. Now, right on cue, Apple has acquired Emotient, a startup that uses artificial intelligence to analyze your facial expressions and read your emotions.
First reported by The Wall Street Journal, the deal is notable because, well, it’s Apple, the world’s most valuable company and one of the most powerful tech giants. It’s unclear how Apple intends to use the company, but as Moore indicates, the tech built by Emotient is part of much larger trend across the industry. Using what are called deep neural networks—vast networks of hardware and software that approximate the web of neurons in the human brain—companies like Google and Facebook are working on similar face recognition technology and have already rolled it into their online services.
'There are huge implications in terms of making dialogue with computers much more meaningful.'
“We have very real data points showing computers doing a better job than humans in accessing emotional states,” Moore said in December. “There are huge implications in terms of making dialogue with computers much more meaningful.”
Moore points our that such technology can be used for everything from security to accessing mental heath. As the Journal explains, Emotient sold its tech to advertisers, letting them analyze how consumers responded to their ads. According to the startup, doctors have also used the technology to determine patient pain, and retailers have used it to track how shoppers react to products in stores.
With deep neural nets, machines can learn to do certain tasks by analyzing large amounts of data. If you feed enough photos of someone smiling into a neural net, for instance, it can learn to understand when someone is happy. And these techniques can be applied to more than just images. They have also proved successful with speech recognition and, to a certain extent, natural language understanding.
Google and Facebook and Microsoft are at the forefront of this deep learning movement. But Apple been pushing in the same direction. In the fall, Apple acquired a startup called VocalIQ, which uses deep neural nets for speech recognition. You might not be able to hide your true feelings from Siri for much longer.
Reference:
http://www.bloomberg.com/news/articles/2016-01-07/apple-buys-startup-that-sees-what-s-behind-your-smile
'via Blog this'
Sunday, January 3, 2016
读Ph.D.,做科研工作的优点 -- 绿卡数据分析 - 未名空间(mitbbs.com)
读Ph.D.,做科研工作的优点 -- 绿卡数据分析 - 未名空间(mitbbs.com): ",来到美国以后取得绿卡的情况是怎么样的呢:按照美国国务院公布的数字,取得
绿卡的人里面,有两万六千左右是在广州办理的亲属移民签证(这里包括美国公民收养
的中国小孩,每年两千多一点)"
'via Blog this'
绿卡的人里面,有两万六千左右是在广州办理的亲属移民签证(这里包括美国公民收养
的中国小孩,每年两千多一点)"
'via Blog this'
Friday, January 1, 2016
Subscribe to:
Posts (Atom)