In the industrial process situation, principal component analysis (PCA) is ageneral method in data reconciliation. However, PCA sometime is unfeasible to nonlinear featureanalysis and limited in application to nonline...In the industrial process situation, principal component analysis (PCA) is ageneral method in data reconciliation. However, PCA sometime is unfeasible to nonlinear featureanalysis and limited in application to nonlinear industrial process. Kernel PCA (KPCA) is extensionof PCA and can be used for nonlinear feature analysis. A nonlinear data reconciliation method basedon KPCA is proposed. The basic idea of this method is that firstly original data are mapped to highdimensional feature space by nonlinear function, and PCA is implemented in the feature space. Thennonlinear feature analysis is implemented and data are reconstructed by using the kernel. The datareconciliation method based on KPCA is applied to ternary distillation column. Simulation resultsshow that this method can filter the noise in measurements of nonlinear process and reconciliateddata can represent the true information of nonlinear process.展开更多
In practical process industries,a variety of online and offline sensors and measuring instruments have been used for process control and monitoring purposes,which indicates that the measurements coming from different ...In practical process industries,a variety of online and offline sensors and measuring instruments have been used for process control and monitoring purposes,which indicates that the measurements coming from different sources are collected at different sampling rates.To build a complete process monitoring strategy,all these multi-rate measurements should be considered for data-based modeling and monitoring.In this paper,a novel kernel multi-rate probabilistic principal component analysis(K-MPPCA)model is proposed to extract the nonlinear correlations among different sampling rates.In the proposed model,the model parameters are calibrated using the kernel trick and the expectation-maximum(EM)algorithm.Also,the corresponding fault detection methods based on the nonlinear features are developed.Finally,a simulated nonlinear case and an actual pre-decarburization unit in the ammonia synthesis process are tested to demonstrate the efficiency of the proposed method.展开更多
Principal Component Analysis(PCA)is one of the most important feature extraction methods,and Kernel Principal Component Analysis(KPCA)is a nonlinear extension of PCA based on kernel methods.In real world,each input da...Principal Component Analysis(PCA)is one of the most important feature extraction methods,and Kernel Principal Component Analysis(KPCA)is a nonlinear extension of PCA based on kernel methods.In real world,each input data may not be fully assigned to one class and it may partially belong to other classes.Based on the theory of fuzzy sets,this paper presents Fuzzy Principal Component Analysis(FPCA)and its nonlinear extension model,i.e.,Kernel-based Fuzzy Principal Component Analysis(KFPCA).The experimental results indicate that the proposed algorithms have good performances.展开更多
Dimensionality reduction techniques play an important role in data mining. Kernel entropy component analysis( KECA) is a newly developed method for data transformation and dimensionality reduction. This paper conducte...Dimensionality reduction techniques play an important role in data mining. Kernel entropy component analysis( KECA) is a newly developed method for data transformation and dimensionality reduction. This paper conducted a comparative study of KECA with other five dimensionality reduction methods,principal component analysis( PCA),kernel PCA( KPCA),locally linear embedding( LLE),laplacian eigenmaps( LAE) and diffusion maps( DM). Three quality assessment criteria, local continuity meta-criterion( LCMC),trustworthiness and continuity measure(T&C),and mean relative rank error( MRRE) are applied as direct performance indexes to assess those dimensionality reduction methods. Moreover,the clustering accuracy is used as an indirect performance index to evaluate the quality of the representative data gotten by those methods. The comparisons are performed on six datasets and the results are analyzed by Friedman test with the corresponding post-hoc tests. The results indicate that KECA shows an excellent performance in both quality assessment criteria and clustering accuracy assessing.展开更多
Panicle swarm optimization (PSO) is an optimization algorithm based on the swarm intelligent principle. In this paper the modified PSO is applied to a kernel principal component analysis ( KPCA ) for an optimal ke...Panicle swarm optimization (PSO) is an optimization algorithm based on the swarm intelligent principle. In this paper the modified PSO is applied to a kernel principal component analysis ( KPCA ) for an optimal kernel function parameter. We first comprehensively considered within-class scatter and between-class scatter of the sample features. Then, the fitness function of an optimized kernel function parameter is constructed, and the particle swarm optimization algorithm with adaptive acceleration (CPSO) is applied to optimizing it. It is used for gearbox condi- tion recognition, and the result is compared with the recognized results based on principal component analysis (PCA). The results show that KPCA optimized by CPSO can effectively recognize fault conditions of the gearbox by reducing bind set-up of the kernel function parameter, and its results of fault recognition outperform those of PCA. We draw the conclusion that KPCA based on CPSO has an advantage in nonlinear feature extraction of mechanical failure, and is helpful for fault condition recognition of complicated machines.展开更多
The precision of the kernel independent component analysis( KICA) algorithm depends on the type and parameter values of kernel function. Therefore,it's of great significance to study the choice method of KICA'...The precision of the kernel independent component analysis( KICA) algorithm depends on the type and parameter values of kernel function. Therefore,it's of great significance to study the choice method of KICA's kernel parameters for improving its feature dimension reduction result. In this paper, a fitness function was established by use of the ideal of Fisher discrimination function firstly. Then the global optimal solution of fitness function was searched by particle swarm optimization( PSO) algorithm and a multi-state information dimension reduction algorithm based on PSO-KICA was established. Finally,the validity of this algorithm to enhance the precision of feature dimension reduction has been proven.展开更多
A novel nonlinear combination process monitoring method was proposed based on techniques with memory effect (multivariate exponentially weighted moving average (MEWMA)) and kernel independent component analysis (KICA)...A novel nonlinear combination process monitoring method was proposed based on techniques with memory effect (multivariate exponentially weighted moving average (MEWMA)) and kernel independent component analysis (KICA). The method was developed for dealing with nonlinear issues and detecting small or moderate drifts in one or more process variables with autocorrelation. MEWMA charts use additional information from the past history of the process for keeping the memory effect of the process behavior trend. KICA is a recently developed statistical technique for revealing hidden, nonlinear statistically independent factors that underlie sets of measurements and it is a two-phase algorithm: whitened kernel principal component analysis (KPCA) plus independent component analysis (ICA). The application to the fluid catalytic cracking unit (FCCU) simulated process indicates that the proposed combined method based on MEWMA and KICA can effectively capture the nonlinear relationship and detect small drifts in process variables. Its performance significantly outperforms monitoring method based on ICA, MEWMA-ICA and KICA, especially for long-term performance deterioration.展开更多
The changes of kernel nutritive components and seed vigor in F1 seeds of sh2 sweet corn during seed development stage were investigated and the relationships between them were analyzed by time series regression (TSR) ...The changes of kernel nutritive components and seed vigor in F1 seeds of sh2 sweet corn during seed development stage were investigated and the relationships between them were analyzed by time series regression (TSR) analysis. The results show that total soluble sugar and reducing sugar contents gradually declined, while starch and soluble protein contents increased throughout the seed development stages. Germination percentage, energy of germination, germination index and vigor index gradually increased along with seed development and reached the highest levels at 38 d after pollination (DAP). The TSR showed that, during 14 to 42 DAP, total soluble sugar content was independent of the vigor parameters determined in present experiment, while the reducing sugar content had a significant effect on seed vigor. TSR equations between seed reducing sugar and seed vigor were also developed. There were negative correlations between the seed reducing sugar content and the germination percentage, energy of germination, germination index and vigor index, respectively. It is suggested that the seed germination, energy of germination, germination index and vigor index could be predicted by the content of reducing sugar in sweet corn seeds during seed development stages.展开更多
Kernal factor analysis (KFA) with vafimax was proposed by using Mercer kernel function which can map the data in the original space to a high-dimensional feature space, and was compared with the kernel principle com...Kernal factor analysis (KFA) with vafimax was proposed by using Mercer kernel function which can map the data in the original space to a high-dimensional feature space, and was compared with the kernel principle component analysis (KPCA). The results show that the best error rate in handwritten digit recognition by kernel factor analysis with vadmax (4.2%) was superior to KPCA (4.4%). The KFA with varimax could more accurately image handwritten digit recognition.展开更多
Currently,some fault prognosis technology occasionally has relatively unsatisfied performance especially for incipient faults in nonlinear processes duo to their large time delay and complex internal connection.To ove...Currently,some fault prognosis technology occasionally has relatively unsatisfied performance especially for incipient faults in nonlinear processes duo to their large time delay and complex internal connection.To overcome this deficiency,multivariate time delay analysis is incorporated into the high sensitive local kernel principal component analysis.In this approach,mutual information estimation and Bayesian information criterion(BIC) are separately used to acquire the correlation degree and time delay of the process variables.Moreover,in order to achieve prediction,time series prediction by back propagation(BP) network is applied whose input is multivariate correlated time series other than the original time series.Then the multivariate time delayed series and future values obtained by time series prediction are combined to construct the input of local kernel principal component analysis(LKPCA) model for incipient fault prognosis.The new method has been exemplified in a simple nonlinear process and the complicated Tennessee Eastman(TE) benchmark process.The results indicate that the new method has superiority in the fault prognosis sensitivity over other traditional fault prognosis methods.展开更多
How to fit a properly nonlinear classification model from conventional well logs to lithofacies is a key problem for machine learning methods.Kernel methods(e.g.,KFD,SVM,MSVM)are effective attempts to solve this issue...How to fit a properly nonlinear classification model from conventional well logs to lithofacies is a key problem for machine learning methods.Kernel methods(e.g.,KFD,SVM,MSVM)are effective attempts to solve this issue due to abilities of handling nonlinear features by kernel functions.Deep mining of log features indicating lithofacies still needs to be improved for kernel methods.Hence,this work employs deep neural networks to enhance the kernel principal component analysis(KPCA)method and proposes a deep kernel method(DKM)for lithofacies identification using well logs.DKM includes a feature extractor and a classifier.The feature extractor consists of a series of KPCA models arranged according to residual network structure.A gradient-free optimization method is introduced to automatically optimize parameters and structure in DKM,which can avoid complex tuning of parameters in models.To test the validation of the proposed DKM for lithofacies identification,an open-sourced dataset with seven con-ventional logs(GR,CAL,AC,DEN,CNL,LLD,and LLS)and lithofacies labels from the Daniudi Gas Field in China is used.There are eight lithofacies,namely clastic rocks(pebbly,coarse,medium,and fine sand-stone,siltstone,mudstone),coal,and carbonate rocks.The comparisons between DKM and three commonly used kernel methods(KFD,SVM,MSVM)show that(1)DKM(85.7%)outperforms SVM(77%),KFD(79.5%),and MSVM(82.8%)in accuracy of lithofacies identification;(2)DKM is about twice faster than the multi-kernel method(MSVM)with good accuracy.The blind well test in Well D13 indicates that compared with the other three methods DKM improves about 24%in accuracy,35%in precision,41%in recall,and 40%in F1 score,respectively.In general,DKM is an effective method for complex lithofacies identification.This work also discussed the optimal structure and classifier for DKM.Experimental re-sults show that(m_(1),m_(2),O)is the optimal model structure and linear svM is the optimal classifier.(m_(1),m_(2),O)means there are m KPCAs,and then m2 residual units.A workflow to determine an optimal classifier in DKM for lithofacies identification is proposed,too.展开更多
In order to quantitatively analyze air traffic operation complexity,multidimensional metrics were selected based on the operational characteristics of traffic flow.The kernel principal component analysis method was ut...In order to quantitatively analyze air traffic operation complexity,multidimensional metrics were selected based on the operational characteristics of traffic flow.The kernel principal component analysis method was utilized to reduce the dimensionality of metrics,therefore to extract crucial information in the metrics.The hierarchical clustering method was used to analyze the complexity of different airspace.Fourteen sectors of Guangzhou Area Control Center were taken as samples.The operation complexity of traffic situation in each sector was calculated based on real flight radar data.Clustering analysis verified the feasibility and rationality of the method,and provided a reference for airspace operation and management.展开更多
Serial Analysis of Gene Expression (SAGE) is a powerful tool to analyze whole-genome expression profiles. SAGE data, characterized by large quantity and high dimensions, need reducing their dimensions and extract feat...Serial Analysis of Gene Expression (SAGE) is a powerful tool to analyze whole-genome expression profiles. SAGE data, characterized by large quantity and high dimensions, need reducing their dimensions and extract feature to improve the accuracy and efficiency when they are used for pattern recognition and clustering analysis. A Poisson Model-based Kernel (PMK) was proposed based on the Poisson distribution of the SAGE data. Kernel Principle Component Analysis (KPCA) with PMK was proposed and used in feature-extract analysis of mouse retinal SAGE data. The computa-tional results show that this algorithm can extract feature effectively and reduce dimensions of SAGE data.展开更多
Investigation of genetic diversity of geographically distant wheat genotypes is </span><span style="font-family:Verdana;">a </span><span style="font-family:Verdana;">useful ...Investigation of genetic diversity of geographically distant wheat genotypes is </span><span style="font-family:Verdana;">a </span><span style="font-family:Verdana;">useful approach in wheat breeding providing efficient crop varieties. This article presents multivariate cluster and principal component analyses (PCA) of some yield traits of wheat, such as thousand-kernel weight (TKW), grain number, grain yield and plant height. Based on the results, an evaluation of economically valuable attributes by eigenvalues made it possible to determine the components that significantly contribute to the yield of common wheat genotypes. Twenty-five genotypes were grouped into four clusters on the basis of average linkage. The PCA showed four principal components (PC) with eigenvalues ></span><span style="font-family:""> </span><span style="font-family:Verdana;">1, explaining approximately 90.8% of the total variability. According to PC analysis, the variance in the eigenvalues was </span><span style="font-family:Verdana;">the </span><span style="font-family:Verdana;">greatest (4.33) for PC-1, PC-2 (1.86) and PC-3 (1.01). The cluster analysis revealed the classification of 25 accessions into four diverse groups. Averages, standard deviations and variances for clusters based on morpho-physiological traits showed that the maximum average values for grain yield (742.2), biomass (1756.7), grains square meter (18</span><span style="font-family:Verdana;">,</span><span style="font-family:Verdana;">373.7), and grains per spike (45.3) were higher in cluster C compared to other clusters. Cluster D exhibited the maximum thousand-kernel weight (TKW) (46.6).展开更多
Unmanned Aerial Vehicles(UAVs)are widely used and meet many demands in military and civilian fields.With the continuous enrichment and extensive expansion of application scenarios,the safety of UAVs is constantly bein...Unmanned Aerial Vehicles(UAVs)are widely used and meet many demands in military and civilian fields.With the continuous enrichment and extensive expansion of application scenarios,the safety of UAVs is constantly being challenged.To address this challenge,we propose algorithms to detect anomalous data collected from drones to improve drone safety.We deployed a one-class kernel extreme learning machine(OCKELM)to detect anomalies in drone data.By default,OCKELM uses the radial basis(RBF)kernel function as the kernel function of themodel.To improve the performance ofOCKELM,we choose a TriangularGlobalAlignmentKernel(TGAK)instead of anRBF Kernel and introduce the Fast Independent Component Analysis(FastICA)algorithm to reconstruct UAV data.Based on the above improvements,we create a novel anomaly detection strategy FastICA-TGAK-OCELM.The method is finally validated on the UCI dataset and detected on the Aeronautical Laboratory Failures and Anomalies(ALFA)dataset.The experimental results show that compared with other methods,the accuracy of this method is improved by more than 30%,and point anomalies are effectively detected.展开更多
基金This project is supported by Special Foundation for Major State Basic Research of China (Project 973, No.G1998030415)
文摘In the industrial process situation, principal component analysis (PCA) is ageneral method in data reconciliation. However, PCA sometime is unfeasible to nonlinear featureanalysis and limited in application to nonlinear industrial process. Kernel PCA (KPCA) is extensionof PCA and can be used for nonlinear feature analysis. A nonlinear data reconciliation method basedon KPCA is proposed. The basic idea of this method is that firstly original data are mapped to highdimensional feature space by nonlinear function, and PCA is implemented in the feature space. Thennonlinear feature analysis is implemented and data are reconstructed by using the kernel. The datareconciliation method based on KPCA is applied to ternary distillation column. Simulation resultsshow that this method can filter the noise in measurements of nonlinear process and reconciliateddata can represent the true information of nonlinear process.
基金supported by Zhejiang Provincial Natural Science Foundation of China(LY19F030003)Key Research and Development Project of Zhejiang Province(2021C04030)+1 种基金the National Natural Science Foundation of China(62003306)Educational Commission Research Program of Zhejiang Province(Y202044842)。
文摘In practical process industries,a variety of online and offline sensors and measuring instruments have been used for process control and monitoring purposes,which indicates that the measurements coming from different sources are collected at different sampling rates.To build a complete process monitoring strategy,all these multi-rate measurements should be considered for data-based modeling and monitoring.In this paper,a novel kernel multi-rate probabilistic principal component analysis(K-MPPCA)model is proposed to extract the nonlinear correlations among different sampling rates.In the proposed model,the model parameters are calibrated using the kernel trick and the expectation-maximum(EM)algorithm.Also,the corresponding fault detection methods based on the nonlinear features are developed.Finally,a simulated nonlinear case and an actual pre-decarburization unit in the ammonia synthesis process are tested to demonstrate the efficiency of the proposed method.
文摘Principal Component Analysis(PCA)is one of the most important feature extraction methods,and Kernel Principal Component Analysis(KPCA)is a nonlinear extension of PCA based on kernel methods.In real world,each input data may not be fully assigned to one class and it may partially belong to other classes.Based on the theory of fuzzy sets,this paper presents Fuzzy Principal Component Analysis(FPCA)and its nonlinear extension model,i.e.,Kernel-based Fuzzy Principal Component Analysis(KFPCA).The experimental results indicate that the proposed algorithms have good performances.
基金Climbing Peak Discipline Project of Shanghai Dianji University,China(No.15DFXK02)Hi-Tech Research and Development Programs of China(No.2007AA041600)
文摘Dimensionality reduction techniques play an important role in data mining. Kernel entropy component analysis( KECA) is a newly developed method for data transformation and dimensionality reduction. This paper conducted a comparative study of KECA with other five dimensionality reduction methods,principal component analysis( PCA),kernel PCA( KPCA),locally linear embedding( LLE),laplacian eigenmaps( LAE) and diffusion maps( DM). Three quality assessment criteria, local continuity meta-criterion( LCMC),trustworthiness and continuity measure(T&C),and mean relative rank error( MRRE) are applied as direct performance indexes to assess those dimensionality reduction methods. Moreover,the clustering accuracy is used as an indirect performance index to evaluate the quality of the representative data gotten by those methods. The comparisons are performed on six datasets and the results are analyzed by Friedman test with the corresponding post-hoc tests. The results indicate that KECA shows an excellent performance in both quality assessment criteria and clustering accuracy assessing.
基金supported by National Natural Science Foundation under Grant No.50875247Shanxi Province Natural Science Foundation under Grant No.2009011026-1
文摘Panicle swarm optimization (PSO) is an optimization algorithm based on the swarm intelligent principle. In this paper the modified PSO is applied to a kernel principal component analysis ( KPCA ) for an optimal kernel function parameter. We first comprehensively considered within-class scatter and between-class scatter of the sample features. Then, the fitness function of an optimized kernel function parameter is constructed, and the particle swarm optimization algorithm with adaptive acceleration (CPSO) is applied to optimizing it. It is used for gearbox condi- tion recognition, and the result is compared with the recognized results based on principal component analysis (PCA). The results show that KPCA optimized by CPSO can effectively recognize fault conditions of the gearbox by reducing bind set-up of the kernel function parameter, and its results of fault recognition outperform those of PCA. We draw the conclusion that KPCA based on CPSO has an advantage in nonlinear feature extraction of mechanical failure, and is helpful for fault condition recognition of complicated machines.
文摘The precision of the kernel independent component analysis( KICA) algorithm depends on the type and parameter values of kernel function. Therefore,it's of great significance to study the choice method of KICA's kernel parameters for improving its feature dimension reduction result. In this paper, a fitness function was established by use of the ideal of Fisher discrimination function firstly. Then the global optimal solution of fitness function was searched by particle swarm optimization( PSO) algorithm and a multi-state information dimension reduction algorithm based on PSO-KICA was established. Finally,the validity of this algorithm to enhance the precision of feature dimension reduction has been proven.
基金The National Natural Science Foundation ofChina(No60504033)
文摘A novel nonlinear combination process monitoring method was proposed based on techniques with memory effect (multivariate exponentially weighted moving average (MEWMA)) and kernel independent component analysis (KICA). The method was developed for dealing with nonlinear issues and detecting small or moderate drifts in one or more process variables with autocorrelation. MEWMA charts use additional information from the past history of the process for keeping the memory effect of the process behavior trend. KICA is a recently developed statistical technique for revealing hidden, nonlinear statistically independent factors that underlie sets of measurements and it is a two-phase algorithm: whitened kernel principal component analysis (KPCA) plus independent component analysis (ICA). The application to the fluid catalytic cracking unit (FCCU) simulated process indicates that the proposed combined method based on MEWMA and KICA can effectively capture the nonlinear relationship and detect small drifts in process variables. Its performance significantly outperforms monitoring method based on ICA, MEWMA-ICA and KICA, especially for long-term performance deterioration.
基金Supported by the 973 project of China (2013CB733600), the National Natural Science Foundation (21176073), the Doctoral Fund of Ministry of Education (20090074110005), the New Century Excellent Talents in University (NCET-09-0346), "Shu Guang" project (09SG29) and the Fundamental Research Funds for the Central Universities.
基金supported by the National Natural Science Foundation of China (No. 30370911)Education Department of Zhejiang Prov-ince, China (No. 20070147)
文摘The changes of kernel nutritive components and seed vigor in F1 seeds of sh2 sweet corn during seed development stage were investigated and the relationships between them were analyzed by time series regression (TSR) analysis. The results show that total soluble sugar and reducing sugar contents gradually declined, while starch and soluble protein contents increased throughout the seed development stages. Germination percentage, energy of germination, germination index and vigor index gradually increased along with seed development and reached the highest levels at 38 d after pollination (DAP). The TSR showed that, during 14 to 42 DAP, total soluble sugar content was independent of the vigor parameters determined in present experiment, while the reducing sugar content had a significant effect on seed vigor. TSR equations between seed reducing sugar and seed vigor were also developed. There were negative correlations between the seed reducing sugar content and the germination percentage, energy of germination, germination index and vigor index, respectively. It is suggested that the seed germination, energy of germination, germination index and vigor index could be predicted by the content of reducing sugar in sweet corn seeds during seed development stages.
基金The National Defence Foundation of China (No.NEWL51435Qt220401)
文摘Kernal factor analysis (KFA) with vafimax was proposed by using Mercer kernel function which can map the data in the original space to a high-dimensional feature space, and was compared with the kernel principle component analysis (KPCA). The results show that the best error rate in handwritten digit recognition by kernel factor analysis with vadmax (4.2%) was superior to KPCA (4.4%). The KFA with varimax could more accurately image handwritten digit recognition.
基金Supported by the National Natural Science Foundation of China(61573051,61472021)the Natural Science Foundation of Beijing(4142039)+1 种基金Open Fund of the State Key Laboratory of Software Development Environment(SKLSDE-2015KF-01)Fundamental Research Funds for the Central Universities(PT1613-05)
文摘Currently,some fault prognosis technology occasionally has relatively unsatisfied performance especially for incipient faults in nonlinear processes duo to their large time delay and complex internal connection.To overcome this deficiency,multivariate time delay analysis is incorporated into the high sensitive local kernel principal component analysis.In this approach,mutual information estimation and Bayesian information criterion(BIC) are separately used to acquire the correlation degree and time delay of the process variables.Moreover,in order to achieve prediction,time series prediction by back propagation(BP) network is applied whose input is multivariate correlated time series other than the original time series.Then the multivariate time delayed series and future values obtained by time series prediction are combined to construct the input of local kernel principal component analysis(LKPCA) model for incipient fault prognosis.The new method has been exemplified in a simple nonlinear process and the complicated Tennessee Eastman(TE) benchmark process.The results indicate that the new method has superiority in the fault prognosis sensitivity over other traditional fault prognosis methods.
基金supported by the National Natural Science Foundation of China(Grant No.42002134)China Postdoctoral Science Foundation(Grant No.2021T140735)Science Foundation of China University of Petroleum,Beijing(Grant Nos.2462020XKJS02 and 2462020YXZZ004).
文摘How to fit a properly nonlinear classification model from conventional well logs to lithofacies is a key problem for machine learning methods.Kernel methods(e.g.,KFD,SVM,MSVM)are effective attempts to solve this issue due to abilities of handling nonlinear features by kernel functions.Deep mining of log features indicating lithofacies still needs to be improved for kernel methods.Hence,this work employs deep neural networks to enhance the kernel principal component analysis(KPCA)method and proposes a deep kernel method(DKM)for lithofacies identification using well logs.DKM includes a feature extractor and a classifier.The feature extractor consists of a series of KPCA models arranged according to residual network structure.A gradient-free optimization method is introduced to automatically optimize parameters and structure in DKM,which can avoid complex tuning of parameters in models.To test the validation of the proposed DKM for lithofacies identification,an open-sourced dataset with seven con-ventional logs(GR,CAL,AC,DEN,CNL,LLD,and LLS)and lithofacies labels from the Daniudi Gas Field in China is used.There are eight lithofacies,namely clastic rocks(pebbly,coarse,medium,and fine sand-stone,siltstone,mudstone),coal,and carbonate rocks.The comparisons between DKM and three commonly used kernel methods(KFD,SVM,MSVM)show that(1)DKM(85.7%)outperforms SVM(77%),KFD(79.5%),and MSVM(82.8%)in accuracy of lithofacies identification;(2)DKM is about twice faster than the multi-kernel method(MSVM)with good accuracy.The blind well test in Well D13 indicates that compared with the other three methods DKM improves about 24%in accuracy,35%in precision,41%in recall,and 40%in F1 score,respectively.In general,DKM is an effective method for complex lithofacies identification.This work also discussed the optimal structure and classifier for DKM.Experimental re-sults show that(m_(1),m_(2),O)is the optimal model structure and linear svM is the optimal classifier.(m_(1),m_(2),O)means there are m KPCAs,and then m2 residual units.A workflow to determine an optimal classifier in DKM for lithofacies identification is proposed,too.
基金co-supported by the National Natural Science Foundation of China(No.61304190)the Fundamental Research Funds for the Central Universities of China(No.NJ20150030)the Youth Science and Technology Innovation Fund(No.NS2014067)
文摘In order to quantitatively analyze air traffic operation complexity,multidimensional metrics were selected based on the operational characteristics of traffic flow.The kernel principal component analysis method was utilized to reduce the dimensionality of metrics,therefore to extract crucial information in the metrics.The hierarchical clustering method was used to analyze the complexity of different airspace.Fourteen sectors of Guangzhou Area Control Center were taken as samples.The operation complexity of traffic situation in each sector was calculated based on real flight radar data.Clustering analysis verified the feasibility and rationality of the method,and provided a reference for airspace operation and management.
基金Supported by the National Natural Science Foundation of China (No. 50877004)
文摘Serial Analysis of Gene Expression (SAGE) is a powerful tool to analyze whole-genome expression profiles. SAGE data, characterized by large quantity and high dimensions, need reducing their dimensions and extract feature to improve the accuracy and efficiency when they are used for pattern recognition and clustering analysis. A Poisson Model-based Kernel (PMK) was proposed based on the Poisson distribution of the SAGE data. Kernel Principle Component Analysis (KPCA) with PMK was proposed and used in feature-extract analysis of mouse retinal SAGE data. The computa-tional results show that this algorithm can extract feature effectively and reduce dimensions of SAGE data.
文摘Investigation of genetic diversity of geographically distant wheat genotypes is </span><span style="font-family:Verdana;">a </span><span style="font-family:Verdana;">useful approach in wheat breeding providing efficient crop varieties. This article presents multivariate cluster and principal component analyses (PCA) of some yield traits of wheat, such as thousand-kernel weight (TKW), grain number, grain yield and plant height. Based on the results, an evaluation of economically valuable attributes by eigenvalues made it possible to determine the components that significantly contribute to the yield of common wheat genotypes. Twenty-five genotypes were grouped into four clusters on the basis of average linkage. The PCA showed four principal components (PC) with eigenvalues ></span><span style="font-family:""> </span><span style="font-family:Verdana;">1, explaining approximately 90.8% of the total variability. According to PC analysis, the variance in the eigenvalues was </span><span style="font-family:Verdana;">the </span><span style="font-family:Verdana;">greatest (4.33) for PC-1, PC-2 (1.86) and PC-3 (1.01). The cluster analysis revealed the classification of 25 accessions into four diverse groups. Averages, standard deviations and variances for clusters based on morpho-physiological traits showed that the maximum average values for grain yield (742.2), biomass (1756.7), grains square meter (18</span><span style="font-family:Verdana;">,</span><span style="font-family:Verdana;">373.7), and grains per spike (45.3) were higher in cluster C compared to other clusters. Cluster D exhibited the maximum thousand-kernel weight (TKW) (46.6).
基金supported by the Natural Science Foundation of The Jiangsu Higher Education Institutions of China(Grant No.19JKB520031).
文摘Unmanned Aerial Vehicles(UAVs)are widely used and meet many demands in military and civilian fields.With the continuous enrichment and extensive expansion of application scenarios,the safety of UAVs is constantly being challenged.To address this challenge,we propose algorithms to detect anomalous data collected from drones to improve drone safety.We deployed a one-class kernel extreme learning machine(OCKELM)to detect anomalies in drone data.By default,OCKELM uses the radial basis(RBF)kernel function as the kernel function of themodel.To improve the performance ofOCKELM,we choose a TriangularGlobalAlignmentKernel(TGAK)instead of anRBF Kernel and introduce the Fast Independent Component Analysis(FastICA)algorithm to reconstruct UAV data.Based on the above improvements,we create a novel anomaly detection strategy FastICA-TGAK-OCELM.The method is finally validated on the UCI dataset and detected on the Aeronautical Laboratory Failures and Anomalies(ALFA)dataset.The experimental results show that compared with other methods,the accuracy of this method is improved by more than 30%,and point anomalies are effectively detected.