用于提高谷歌图像搜索结果的二分类器在线学习方法(英文)

来源 :自动化学报 | 被引量 : 0次 | 上传用户:ganmaogaishilangren
下载到本地 , 更方便阅读
声明 : 本文档内容版权归属内容提供方 , 如果您对本文有版权争议 , 可与客服联系进行内容授权或下架
论文部分内容阅读
It is promising to improve web image search results through exploiting the results visual contents for learning a binary classifier which is used to refine the results relevance degrees to the given query. This paper proposes an algorithm framework as a solution to this problem and investigates the key issue of training data selection under the framework. The training data selection process is divided into two stages: initial selection for triggering the classifier learning and dynamic selection in the iterations of classifier learning. We investigate two main ways of initial training data selection, including clustering based and ranking based, and compare automatic training data selection schemes with manual manner. Furthermore, support vector machines and the max-min pseudo-probability(MMP) based Bayesian classifier are employed to support image classification, respectively. By varying these factors in the framework, we implement eight algorithms and tested them on keyword based image search results from Google search engine. The experimental results confirm that how to select the training data from noisy search results is really a key issue in the problem considered in this paper and show that the proposed algorithm is effective to improve Google search results, especially at top ranks, thus is helpful to reduce the user labor in finding the desired images by browsing the ranking in depth. Even so, it is still worth meditative to make automatic training data selection scheme better towards perfect human annotation. It is promising to improve web image search results through exploiting the results visual contents for learning a binary classifier which is used to refine the results relevance degrees to the given query. This paper proposes an algorithm framework as a solution to this problem and investigates the key issue of training data selection under the framework. The training data selection process is divided into two stages: initial selection for triggering the classifier learning and dynamic selection in the iterations of classifier learning. We investigate two main ways of initial training data selection, including clustering Furthermore, support vector machines and the max-min pseudo-probability (MMP) based Bayesian classifier are employed to support image classification, respectively. By varying these factors in the framework, we implement eight algorithms and tested them on keyword based ima ge search results from Google search engine. The experimental results confirm that how to select the training data from noisy search results is really a key issue in the problem considered in this paper and show that the proposed algorithm is effective to improve Google search results, especially at top ranks, thus is helpful to reduce the user labor in finding the desired images by browsing the ranking in depth. Even so, it is still worth meditative to make automatic training data selection scheme better towards perfect human annotation.
其他文献
本论文主要介绍了利用红外天文望远镜的观测和巡天数据研究致密星系统红外辐射的起源和过程,以及致密星系统的形成和演化,同时搜寻致密星尘埃盘。内容包括了对毫秒脉冲星双星的
本文主要是关于国家大科学项目大天区面积多目标光纤光谱望远镜(The Large SkyArea Multi-object Fiber Spectroscopic Telescope,简称LAMOST)输入星表中恒星与星系的分离,使用
我国陶瓷文化渊源流长,陶瓷灯具是自古以来的陶瓷品种之一,从古代的油灯、烛台,到现代社会的陶瓷灯罩,都是生活中实用之器。陶瓷灯具之所以在生活中能够占有一席之地,与它的
在一场展览中,相关设计人员对于展品展示空间所进行的科学合理的设计,对于展览的顺利进行具有十分巨大的影响作用。本文针对展品的整体与局部布局的科学安排、色彩艺术在展示
宇宙的加速膨胀和暗能量是当前宇宙学研究中最热门的前沿问题之一。暗能量的理论模型很多,如宇宙常数人、标量场quintessence、K-essence、phantom、quintom、Chaplygin气体等
本文通过对巴尔蒂斯作品中的构图分析来更加深入的去理解构图,总结出巴尔蒂斯如何通过各种的构图形式来营造画面气氛,达到其独特的艺术效果以及他对后人所带来的启示。 This
天文望远镜技术一直代表着当世科学技术发展的最高水平,大口径光学非球面是其中的关键技术之一。按照现阶段发展趋势,十米及以上级天文望远镜镜面主要采用子镜拼接技术,其子镜为
我们在文中给出了一个理论模型来描述星系核球的形成过程中恒星形成活动与黑洞增长之间的关系。我们发现嵌在背景引力场中的致密气体存在一个三维临界密度,超过这个密度时,气体
本文以2008年沪深股市发行A股的公司为样本,对我国上市公司会计政策组合选择与成本费用的相关性进行了检验。结果发现:会计政策选择与管理费用、公允价值变动损益具有很强的
二十世纪六十年代至九十年代,世界各国相继开展了基于模拟电视体制的高精度时间频率传递技术研究工作,出台了模拟电视授时相关技术的系列标准,成功的将模拟电视授时推广应用于多