{"id":1678,"date":"2018-07-10T17:26:51","date_gmt":"2018-07-10T09:26:51","guid":{"rendered":"https:\/\/yanjingang.com\/blog\/?p=1678"},"modified":"2019-01-30T13:51:46","modified_gmt":"2019-01-30T05:51:46","slug":"%e5%b0%8f%e7%8c%aa%e5%ad%a6paddle-word-embedding%e5%b1%82%e6%98%af%e5%81%9a%e4%bb%80%e4%b9%88%e7%94%a8%e7%9a%84%ef%bc%9f","status":"publish","type":"post","link":"https:\/\/yanjingang.com\/blog\/?p=1678","title":{"rendered":"\u5c0f\u732a\u5b66AI\u2014word embedding\u5c42\u662f\u505a\u4ec0\u4e48\u7528\u7684\uff1f"},"content":{"rendered":"<div>\n<div>\n<p>word embedding\u7684\u610f\u601d\u662f\uff1a\u7ed9\u51fa\u4e00\u4e2a\u6587\u6863\uff0c\u6587\u6863\u5c31\u662f\u4e00\u4e2a\u5355\u8bcd\u5e8f\u5217\u6bd4\u5982 \u201cA B A C B F G\u201d, \u5e0c\u671b\u5bf9\u6587\u6863\u4e2d\u6bcf\u4e2a\u4e0d\u540c\u7684\u5355\u8bcd\u90fd\u5f97\u5230\u4e00\u4e2a\u5bf9\u5e94\u7684\u5411\u91cf(\u5f80\u5f80\u662f\u4f4e\u7ef4\u5411\u91cf)\u8868\u793a\u3002<br \/>\n\u6bd4\u5982\uff0c\u5bf9\u4e8e\u8fd9\u6837\u7684\u201cA B A C B F G\u201d\u7684\u4e00\u4e2a\u5e8f\u5217\uff0c\u4e5f\u8bb8\u6211\u4eec\u6700\u540e\u80fd\u5f97\u5230\uff1aA\u5bf9\u5e94\u7684\u5411\u91cf\u4e3a[0.1 0.6 -0.5]\uff0cB\u5bf9\u5e94\u7684\u5411\u91cf\u4e3a[-0.2 0.9 0.7] \uff08\u6b64\u5904\u7684\u6570\u503c\u53ea\u7528\u4e8e\u793a\u610f\uff09<\/p>\n<p><span style=\"color: #ff0000;\">\u4e4b\u6240\u4ee5\u5e0c\u671b\u628a\u6bcf\u4e2a\u5355\u8bcd\u53d8\u6210\u4e00\u4e2a\u5411\u91cf\uff0c\u76ee\u7684\u8fd8\u662f\u4e3a\u4e86\u65b9\u4fbf\u8ba1\u7b97\uff0c\u6bd4\u5982\u201c\u6c42\u5355\u8bcdA\u7684\u540c\u4e49\u8bcd\u201d\uff0c\u5c31\u53ef\u4ee5\u901a\u8fc7\u201c\u6c42\u4e0e\u5355\u8bcdA\u5728cos\u8ddd\u79bb\u4e0b\u6700\u76f8\u4f3c\u7684\u5411\u91cf\u201d\u6765\u505a\u5230\u3002<\/span><\/p>\n<p>word embedding\u4e0d\u662f\u4e00\u4e2a\u65b0\u7684topic\uff0c\u5f88\u65e9\u5c31\u5df2\u7ecf\u6709\u4eba\u505a\u4e86\uff0c\u6bd4\u5982bengio\u7684paper\u201cNeural probabilistic language models\u201d\uff0c\u8fd9\u5176\u5b9e\u8fd8\u4e0d\u7b97\u6700\u65e9\uff0c\u66f4\u65e9\u7684\u65f6\u5019\uff0cHinton\u5c31\u5df2\u7ecf\u63d0\u51fa\u4e86distributed representation\u7684\u6982\u5ff5\u201cLearning distributed representations of concepts\u201d(\u53ea\u4e0d\u8fc7\u4e0d\u662f\u7528\u5728word embedding\u4e0a\u9762) \uff0cAAAI2015\u7684\u65f6\u5019\u95ee\u8fc7Hinton\u600e\u4e48\u770bgoogle\u7684word2vec\uff0c\u4ed6\u8bf4\u81ea\u5df120\u5e74\u524d\u5c31\u5df2\u7ecf\u641e\u8fc7\u4e86\uff0c\u54c8\u54c8\uff0c\u4f30\u8ba1\u6307\u7684\u5c31\u662f\u8fd9\u7bc7paper\u3002<\/p>\n<p>\u603b\u4e4b\uff0c\u5e38\u89c1\u7684word embedding\u65b9\u6cd5\u5c31\u662f\u5148\u4ece\u6587\u672c\u4e2d\u4e3a\u6bcf\u4e2a\u5355\u8bcd\u6784\u9020\u4e00\u7ec4features\uff0c\u7136\u540e\u5bf9\u8fd9\u7ec4feature\u505adistributed representations\uff0c\u54c8\u54c8\uff0c\u76f8\u6bd4\u4e8e\u4f20\u7edf\u7684distributed representations\uff0c\u533a\u522b\u5c31\u662f\u591a\u4e86\u4e00\u6b65(\u5148\u4ece\u6587\u6863\u4e2d\u4e3a\u6bcf\u4e2a\u5355\u8bcd\u6784\u9020\u4e00\u7ec4feature)\u3002<\/p>\n<p>\u65e2\u7136word embedding\u662f\u4e00\u4e2a\u8001\u7684topic\uff0c\u4e3a\u4ec0\u4e48\u4f1a\u706b\u5462\uff1f\u539f\u56e0\u662fTomas Mikolov\u5728Google\u7684\u65f6\u5019\u53d1\u7684\u8fd9\u4e24\u7bc7paper\uff1a\u201cEfficient Estimation of Word Representations in Vector Space\u201d\u3001\u201cDistributed Representations of Words and Phrases and their Compositionality\u201d\u3002<\/p>\n<p>\u8fd9\u4e24\u7bc7paper\u4e2d\u63d0\u51fa\u4e86\u4e00\u4e2aword2vec\u7684\u5de5\u5177\u5305\uff0c\u91cc\u9762\u5305\u542b\u4e86\u51e0\u79cdword embedding\u7684\u65b9\u6cd5\uff0c\u8fd9\u4e9b\u65b9\u6cd5\u6709\u4e24\u4e2a\u7279\u70b9\u3002\u4e00\u4e2a\u7279\u70b9\u662f\u901f\u5ea6\u5feb\uff0c\u53e6\u4e00\u4e2a\u7279\u70b9\u662f\u5f97\u5230\u7684embedding vectors\u5177\u5907analogy\u6027\u8d28\u3002analogy\u6027\u8d28\u7c7b\u4f3c\u4e8e\u201cA-B=C-D\u201d\u8fd9\u6837\u7684\u7ed3\u6784\uff0c\u4e3e\u4f8b\u8bf4\u660e\uff1a\u201c\u5317\u4eac-\u4e2d\u56fd = \u5df4\u9ece-\u6cd5\u56fd\u201d\u3002Tomas Mikolov\u8ba4\u4e3a\u5177\u5907\u8fd9\u6837\u7684\u6027\u8d28\uff0c\u5219\u8bf4\u660e\u5f97\u5230\u7684embedding vectors\u6027\u8d28\u975e\u5e38\u597d\uff0c\u80fd\u591fmodel\u5230\u8bed\u4e49\u3002<\/p>\n<p>\u8fd9\u4e24\u7bc7paper\u662f2013\u5e74\u7684\u5de5\u4f5c\uff0c\u81f3\u4eca(2015.8)\uff0c\u8fd9\u4e24\u7bc7paper\u7684\u5f15\u7528\u91cf\u65e9\u5df2\u7ecf\u8d85\u597d\u51e0\u767e\uff0c\u8db3\u4ee5\u770b\u51fa\u5176\u5f71\u54cd\u529b\u5f88\u5927\u3002\u5f53\u7136\uff0cword embedding\u7684\u65b9\u6848\u8fd8\u6709\u5f88\u591a\uff0c\u5e38\u89c1\u7684word embedding\u7684\u65b9\u6cd5\u6709:<br \/>\n1. Distributed Representations of Words and Phrases and their Compositionality<br \/>\n2. Efficient Estimation of Word Representations in Vector Space<br \/>\n3. GloVe Global Vectors forWord Representation<br \/>\n4. Neural probabilistic language models<br \/>\n5. Natural language processing (almost) from scratch<br \/>\n6. Learning word embeddings efficiently with noise contrastive estimation<br \/>\n7. A scalable hierarchical distributed language model<br \/>\n8. Three new graphical models for statistical language modelling<br \/>\n9. Improving word representations via global context and multiple word prototypes<\/p>\n<p>word2vec\u4e2d\u7684\u6a21\u578b\u81f3\u4eca(2015.8)\u8fd8\u662f\u5b58\u5728\u4e0d\u5c11\u672a\u89e3\u4e4b\u8c1c\uff0c\u56e0\u6b64\u5c31\u6709\u4e0d\u5c11papers\u5c1d\u8bd5\u53bb\u89e3\u91ca\u5176\u4e2d\u4e00\u4e9b\u8c1c\u56e2\uff0c\u6216\u8005\u5efa\u7acb\u5176\u4e0e\u5176\u4ed6\u6a21\u578b\u4e4b\u95f4\u7684\u8054\u7cfb\uff0c\u4e0b\u9762\u662fpaper list<br \/>\n1. Neural Word Embeddings as Implicit Matrix Factorization<br \/>\n2. Linguistic Regularities in Sparse and Explicit Word Representation<br \/>\n3. Random Walks on Context Spaces Towards an Explanation of the Mysteries of Semantic Word Embeddings<br \/>\n4. word2vec Explained Deriving Mikolov et al.\u2019s Negative Sampling Word Embedding Method<br \/>\n5. Linking GloVe with word2vec<br \/>\n6. Word Embedding Revisited: A New Representation Learning and Explicit Matrix Factorization Perspective<\/p>\n<\/div>\n<\/div>\n<div><\/div>\n<div>\u6458\u81ea\uff1ahttps:\/\/www.zhihu.com\/question\/32275069\/answer\/61059440<\/div>\n","protected":false},"excerpt":{"rendered":"<p>word embedding\u7684\u610f\u601d\u662f\uff1a\u7ed9\u51fa\u4e00\u4e2a\u6587\u6863\uff0c\u6587\u6863\u5c31\u662f\u4e00\u4e2a\u5355\u8bcd\u5e8f\u5217\u6bd4\u5982 \u201cA B A C B F G\u201d, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[535,539],"tags":[579,536,577,578,732,581,580],"_links":{"self":[{"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=\/wp\/v2\/posts\/1678"}],"collection":[{"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1678"}],"version-history":[{"count":0,"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=\/wp\/v2\/posts\/1678\/revisions"}],"wp:attachment":[{"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1678"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1678"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/yanjingang.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1678"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}