As massive underground projects have become popular in dense urban cities,a problem has arisen:which model predicts the best for Tunnel Boring Machine(TBM)performance in these tunneling projects?However,performance le...As massive underground projects have become popular in dense urban cities,a problem has arisen:which model predicts the best for Tunnel Boring Machine(TBM)performance in these tunneling projects?However,performance level of TBMs in complex geological conditions is still a great challenge for practitioners and researchers.On the other hand,a reliable and accurate prediction of TBM performance is essential to planning an applicable tunnel construction schedule.The performance of TBM is very difficult to estimate due to various geotechnical and geological factors and machine specifications.The previously-proposed intelligent techniques in this field are mostly based on a single or base model with a low level of accuracy.Hence,this study aims to introduce a hybrid randomforest(RF)technique optimized by global harmony search with generalized oppositionbased learning(GOGHS)for forecasting TBM advance rate(AR).Optimizing the RF hyper-parameters in terms of,e.g.,tree number and maximum tree depth is the main objective of using the GOGHS-RF model.In the modelling of this study,a comprehensive databasewith themost influential parameters onTBMtogetherwithTBM AR were used as input and output variables,respectively.To examine the capability and power of the GOGHSRF model,three more hybrid models of particle swarm optimization-RF,genetic algorithm-RF and artificial bee colony-RF were also constructed to forecast TBM AR.Evaluation of the developed models was performed by calculating several performance indices,including determination coefficient(R2),root-mean-square-error(RMSE),and mean-absolute-percentage-error(MAPE).The results showed that theGOGHS-RF is a more accurate technique for estimatingTBMAR compared to the other applied models.The newly-developedGOGHS-RFmodel enjoyed R2=0.9937 and 0.9844,respectively,for train and test stages,which are higher than a pre-developed RF.Also,the importance of the input parameters was interpreted through the SHapley Additive exPlanations(SHAP)method,and it was found that thrust force per cutter is the most important variable on TBMAR.The GOGHS-RF model can be used in mechanized tunnel projects for predicting and checking performance.展开更多
Real-time intelligent lithology identification while drilling is vital to realizing downhole closed-loop drilling. The complex and changeable geological environment in the drilling makes lithology identification face ...Real-time intelligent lithology identification while drilling is vital to realizing downhole closed-loop drilling. The complex and changeable geological environment in the drilling makes lithology identification face many challenges. This paper studies the problems of difficult feature information extraction,low precision of thin-layer identification and limited applicability of the model in intelligent lithologic identification. The author tries to improve the comprehensive performance of the lithology identification model from three aspects: data feature extraction, class balance, and model design. A new real-time intelligent lithology identification model of dynamic felling strategy weighted random forest algorithm(DFW-RF) is proposed. According to the feature selection results, gamma ray and 2 MHz phase resistivity are the logging while drilling(LWD) parameters that significantly influence lithology identification. The comprehensive performance of the DFW-RF lithology identification model has been verified in the application of 3 wells in different areas. By comparing the prediction results of five typical lithology identification algorithms, the DFW-RF model has a higher lithology identification accuracy rate and F1 score. This model improves the identification accuracy of thin-layer lithology and is effective and feasible in different geological environments. The DFW-RF model plays a truly efficient role in the realtime intelligent identification of lithologic information in closed-loop drilling and has greater applicability, which is worthy of being widely used in logging interpretation.展开更多
Driven piles are used in many geological environments as a practical and convenient structural component.Hence,the determination of the drivability of piles is actually of great importance in complex geotechnical appl...Driven piles are used in many geological environments as a practical and convenient structural component.Hence,the determination of the drivability of piles is actually of great importance in complex geotechnical applications.Conventional methods of predicting pile drivability often rely on simplified physicalmodels or empirical formulas,whichmay lack accuracy or applicability in complex geological conditions.Therefore,this study presents a practical machine learning approach,namely a Random Forest(RF)optimized by Bayesian Optimization(BO)and Particle Swarm Optimization(PSO),which not only enhances prediction accuracy but also better adapts to varying geological environments to predict the drivability parameters of piles(i.e.,maximumcompressive stress,maximum tensile stress,and blow per foot).In addition,support vector regression,extreme gradient boosting,k nearest neighbor,and decision tree are also used and applied for comparison purposes.In order to train and test these models,among the 4072 datasets collected with 17model inputs,3258 datasets were randomly selected for training,and the remaining 814 datasets were used for model testing.Lastly,the results of these models were compared and evaluated using two performance indices,i.e.,the root mean square error(RMSE)and the coefficient of determination(R2).The results indicate that the optimized RF model achieved lower RMSE than other prediction models in predicting the three parameters,specifically 0.044,0.438,and 0.146;and higher R^(2) values than other implemented techniques,specifically 0.966,0.884,and 0.977.In addition,the sensitivity and uncertainty of the optimized RF model were analyzed using Sobol sensitivity analysis and Monte Carlo(MC)simulation.It can be concluded that the optimized RF model could be used to predict the performance of the pile,and it may provide a useful reference for solving some problems under similar engineering conditions.展开更多
Fatigue reliability-based design optimization of aeroengine structures involves multiple repeated calculations of reliability degree and large-scale calls of implicit high-nonlinearity limit state function,leading to ...Fatigue reliability-based design optimization of aeroengine structures involves multiple repeated calculations of reliability degree and large-scale calls of implicit high-nonlinearity limit state function,leading to the traditional direct Monte Claro and surrogate methods prone to unacceptable computing efficiency and accuracy.In this case,by fusing the random subspace strategy and weight allocation technology into bagging ensemble theory,a random forest(RF)model is presented to enhance the computing efficiency of reliability degree;moreover,by embedding the RF model into multilevel optimization model,an efficient RF-assisted fatigue reliability-based design optimization framework is developed.Regarding the low-cycle fatigue reliability-based design optimization of aeroengine turbine disc as a case,the effectiveness of the presented framework is validated.The reliabilitybased design optimization results exhibit that the proposed framework holds high computing accuracy and computing efficiency.The current efforts shed a light on the theory/method development of reliability-based design optimization of complex engineering structures.展开更多
In the era of the Internet,widely used web applications have become the target of hacker attacks because they contain a large amount of personal information.Among these vulnerabilities,stealing private data through cr...In the era of the Internet,widely used web applications have become the target of hacker attacks because they contain a large amount of personal information.Among these vulnerabilities,stealing private data through crosssite scripting(XSS)attacks is one of the most commonly used attacks by hackers.Currently,deep learning-based XSS attack detection methods have good application prospects;however,they suffer from problems such as being prone to overfitting,a high false alarm rate,and low accuracy.To address these issues,we propose a multi-stage feature extraction and fusion model for XSS detection based on Random Forest feature enhancement.The model utilizes RandomForests to capture the intrinsic structure and patterns of the data by extracting leaf node indices as features,which are subsequentlymergedwith the original data features to forma feature setwith richer information content.Further feature extraction is conducted through three parallel channels.Channel I utilizes parallel onedimensional convolutional layers(1Dconvolutional layers)with different convolutional kernel sizes to extract local features at different scales and performmulti-scale feature fusion;Channel II employsmaximum one-dimensional pooling layers(max 1D pooling layers)of various sizes to extract key features from the data;and Channel III extracts global information bi-directionally using a Bi-Directional Long-Short TermMemory Network(Bi-LSTM)and incorporates a multi-head attention mechanism to enhance global features.Finally,effective classification and prediction of XSS are performed by fusing the features of the three channels.To test the effectiveness of the model,we conduct experiments on six datasets.We achieve an accuracy of 100%on the UNSW-NB15 dataset and 99.99%on the CICIDS2017 dataset,which is higher than that of the existing models.展开更多
Automatically detecting Ulva prolifera(U.prolifera)in rainy and cloudy weather using remote sensing imagery has been a long-standing problem.Here,we address this challenge by combining high-resolution Synthetic Apertu...Automatically detecting Ulva prolifera(U.prolifera)in rainy and cloudy weather using remote sensing imagery has been a long-standing problem.Here,we address this challenge by combining high-resolution Synthetic Aperture Radar(SAR)imagery with the machine learning,and detect the U.prolifera of the South Yellow Sea of China(SYS)in 2021.The findings indicate that the Random Forest model can accurately and robustly detect U.prolifera,even in the presence of complex ocean backgrounds and speckle noise.Visual inspection confirmed that the method successfully identified the majority of pixels containing U.prolifera without misidentify-ing noise pixels or seawater pixels as U.prolifera.Additionally,the method demonstrated consistent performance across different im-ages,with an average Area Under Curve(AUC)of 0.930(+0.028).The analysis yielded an overall accuracy of over 96%,with an aver-age Kappa coefficient of 0.941(+0.038).Compared to the traditional thresholding method,Random Forest model has a lower estima-tion error of 14.81%.Practical application indicates that this method can be used in the detection of unprecedented U.prolifera in 2021 to derive continuous spatiotemporal changes.This study provides a potential new method to detect U.prolifera and enhances our under-standing of macroalgal outbreaks in the marine environment.展开更多
Precise and timely prediction of crop yields is crucial for food security and the development of agricultural policies.However,crop yield is influenced by multiple factors within complex growth environments.Previous r...Precise and timely prediction of crop yields is crucial for food security and the development of agricultural policies.However,crop yield is influenced by multiple factors within complex growth environments.Previous research has paid relatively little attention to the interference of environmental factors and drought on the growth of winter wheat.Therefore,there is an urgent need for more effective methods to explore the inherent relationship between these factors and crop yield,making precise yield prediction increasingly important.This study was based on four type of indicators including meteorological,crop growth status,environmental,and drought index,from October 2003 to June 2019 in Henan Province as the basic data for predicting winter wheat yield.Using the sparrow search al-gorithm combined with random forest(SSA-RF)under different input indicators,accuracy of winter wheat yield estimation was calcu-lated.The estimation accuracy of SSA-RF was compared with partial least squares regression(PLSR),extreme gradient boosting(XG-Boost),and random forest(RF)models.Finally,the determined optimal yield estimation method was used to predict winter wheat yield in three typical years.Following are the findings:1)the SSA-RF demonstrates superior performance in estimating winter wheat yield compared to other algorithms.The best yield estimation method is achieved by four types indicators’composition with SSA-RF)(R^(2)=0.805,RRMSE=9.9%.2)Crops growth status and environmental indicators play significant roles in wheat yield estimation,accounting for 46%and 22%of the yield importance among all indicators,respectively.3)Selecting indicators from October to April of the follow-ing year yielded the highest accuracy in winter wheat yield estimation,with an R^(2)of 0.826 and an RMSE of 9.0%.Yield estimates can be completed two months before the winter wheat harvest in June.4)The predicted performance will be slightly affected by severe drought.Compared with severe drought year(2011)(R^(2)=0.680)and normal year(2017)(R^(2)=0.790),the SSA-RF model has higher prediction accuracy for wet year(2018)(R^(2)=0.820).This study could provide an innovative approach for remote sensing estimation of winter wheat yield.yield.展开更多
This study proposed a new real-time manufacturing process monitoring method to monitor and detect process shifts in manufacturing operations.Since real-time production process monitoring is critical in today’s smart ...This study proposed a new real-time manufacturing process monitoring method to monitor and detect process shifts in manufacturing operations.Since real-time production process monitoring is critical in today’s smart manufacturing.The more robust the monitoring model,the more reliable a process is to be under control.In the past,many researchers have developed real-time monitoring methods to detect process shifts early.However,thesemethods have limitations in detecting process shifts as quickly as possible and handling various data volumes and varieties.In this paper,a robust monitoring model combining Gated Recurrent Unit(GRU)and Random Forest(RF)with Real-Time Contrast(RTC)called GRU-RF-RTC was proposed to detect process shifts rapidly.The effectiveness of the proposed GRU-RF-RTC model is first evaluated using multivariate normal and nonnormal distribution datasets.Then,to prove the applicability of the proposed model in a realmanufacturing setting,the model was evaluated using real-world normal and non-normal problems.The results demonstrate that the proposed GRU-RF-RTC outperforms other methods in detecting process shifts quickly with the lowest average out-of-control run length(ARL1)in all synthesis and real-world problems under normal and non-normal cases.The experiment results on real-world problems highlight the significance of the proposed GRU-RF-RTC model in modern manufacturing process monitoring applications.The result reveals that the proposed method improves the shift detection capability by 42.14%in normal and 43.64%in gamma distribution problems.展开更多
Information on the decay process of nuclides in the superheavy region is critical in investigating new elements beyond oganesson and the island of stability.This paper presents the application of a random forest algor...Information on the decay process of nuclides in the superheavy region is critical in investigating new elements beyond oganesson and the island of stability.This paper presents the application of a random forest algorithm to examine the competition among different decay modes in the superheavy region,includingα decay,β^(-)decay,β^(+)decay,electron capture and spontaneous fission.The observed half-lives and dominant decay mode are well reproduced.The dominant decay mode of 96.9%of the nuclei beyond ^(212) Po is correctly obtained.Further,α decay is predicted to be the dominant decay mode for isotopes in new elements Z=119-122,except for spontaneous fission in certain even–even elements owing to the increased Coulomb repulsion and odd–even effect.The predicted half-lives demonstrate the existence of a long-lived spontaneous fission island southwest of ^(298) Fl caused by the competition between the fission barrier and Coulomb repulsion.A better understanding of spontaneous fission,particularly beyond ^(286)Fl,is crucial in the search for new elements and the island of stability.展开更多
Acid production with flue gas is a complex nonlinear process with multiple variables and strong coupling.The operation data is an important basis for state monitoring,optimal control,and fault diagnosis.However,the op...Acid production with flue gas is a complex nonlinear process with multiple variables and strong coupling.The operation data is an important basis for state monitoring,optimal control,and fault diagnosis.However,the operating environment of acid production with flue gas is complex and there is much equipment.The data obtained by the detection equipment is seriously polluted and prone to abnormal phenomena such as data loss and outliers.Therefore,to solve the problem of abnormal data in the process of acid production with flue gas,a data cleaning method based on improved random forest is proposed.Firstly,an outlier data recognition model based on isolation forest is designed to identify and eliminate the outliers in the dataset.Secondly,an improved random forest regression model is established.Genetic algorithm is used to optimize the hyperparameters of the random forest regression model.Then the optimal parameter combination is found in the search space and the trend of data is predicted.Finally,the improved random forest data cleaning method is used to compensate for the missing data after eliminating abnormal data and the data cleaning is realized.Results show that the proposed method can accurately eliminate and compensate for the abnormal data in the process of acid production with flue gas.The method improves the accuracy of compensation for missing data.With the data after cleaning,a more accurate model can be established,which is significant to the subsequent temperature control.The conversion rate of SO_(2) can be further improved,thereby improving the yield of sulfuric acid and economic benefits.展开更多
60 GHz millimeter wave(mmWave)system provides extremely high time resolution and multipath components(MPC)separation and has great potential to achieve high precision in the indoor positioning.However,the ranging data...60 GHz millimeter wave(mmWave)system provides extremely high time resolution and multipath components(MPC)separation and has great potential to achieve high precision in the indoor positioning.However,the ranging data is often contaminated by non-line-of-sight(NLOS)transmission.First,six features of 60GHz mm Wave signal under LOS and NLOS conditions are evaluated.Next,a classifier constructed by random forest(RF)algorithm is used to identify line-of-sight(LOS)or NLOS channel.The identification mechanism has excellent generalization performance and the classification accuracy is over 97%.Finally,based on the identification results,a residual weighted least squares positioning method is proposed.All ranging information including that under NLOS channels is fully utilized,positioning failure caused by insufficient LOS links can be avoided.Compared with the conventional least squares approach,the positioning error of the proposed algorithm is reduced by 49%.展开更多
AIM:To analyze ultrasound biomicroscopy(UBM)images using random forest network to find new features to make predictions about vault after implantable collamer lens(ICL)implantation.METHODS:A total of 450 UBM images we...AIM:To analyze ultrasound biomicroscopy(UBM)images using random forest network to find new features to make predictions about vault after implantable collamer lens(ICL)implantation.METHODS:A total of 450 UBM images were collected from the Lixiang Eye Hospital to provide the patient’s preoperative parameters as well as the vault of the ICL after implantation.The vault was set as the prediction target,and the input elements were mainly ciliary sulcus shape parameters,which included 6 angular parameters,2 area parameters,and 2 parameters,distance between ciliary sulci,and anterior chamber height.A random forest regression model was applied to predict the vault,with the number of base estimators(n_estimators)of 2000,the maximum tree depth(max_depth)of 17,the number of tree features(max_features)of Auto,and the random state(random_state)of 40.0.RESULTS:Among the parameters selected in this study,the distance between ciliary sulci had a greater importance proportion,reaching 52%before parameter optimization is performed,and other features had less influence,with an importance proportion of about 5%.The importance of the distance between the ciliary sulci increased to 53% after parameter optimization,and the importance of angle 3 and area 1 increased to 5% and 8%respectively,while the importance of the other parameters remained unchanged,and the distance between the ciliary sulci was considered the most important feature.Other features,although they accounted for a relatively small proportion,also had an impact on the vault prediction.After parameter optimization,the best prediction results were obtained,with a predicted mean value of 763.688μm and an actual mean value of 776.9304μm.The R²was 0.4456 and the root mean square error was 201.5166.CONCLUSION:A study based on UBM images using random forest network can be performed for prediction of the vault after ICL implantation and can provide some reference for ICL size selection.展开更多
A huge number of old arch bridges located in rural regions are at the peak of maintenance.The health monitoring technology of the long-span bridge is hardly applicable to the small-span bridge,owing to the absence of ...A huge number of old arch bridges located in rural regions are at the peak of maintenance.The health monitoring technology of the long-span bridge is hardly applicable to the small-span bridge,owing to the absence of technical resources and sufficient funds in rural regions.There is an urgent need for an economical,fast,and accurate damage identification solution.The authors proposed a damage identification system of an old arch bridge implemented with amachine learning algorithm,which took the vehicle-induced response as the excitation.A damage index was defined based on wavelet packet theory,and a machine learning sample database collecting the denoised response was constructed.Through comparing three machine learning algorithms:Back-Propagation Neural Network(BPNN),Support Vector Machine(SVM),and Random Forest(R.F.),the R.F.damage identification model were found to have a better recognition ability.Finally,the Particle Swarm Optimization(PSO)algorithm was used to optimize the number of subtrees and split features of the R.F.model.The PSO optimized R.F.model was capable of the identification of different damage levels of old arch bridges with sensitive damage index.The proposed framework is practical and promising for the old bridge’s structural damage identification in rural regions.展开更多
As a complex hot problem in the financial field,stock trend forecasting uses a large amount of data and many related indicators;hence it is difficult to obtain sustainable and effective results only by relying on empi...As a complex hot problem in the financial field,stock trend forecasting uses a large amount of data and many related indicators;hence it is difficult to obtain sustainable and effective results only by relying on empirical analysis.Researchers in the field of machine learning have proved that random forest can form better judgements on this kind of problem,and it has an auxiliary role in the prediction of stock trend.This study uses historical trading data of four listed companies in the USA stock market,and the purpose of this study is to improve the performance of random forest model in medium-and long-term stock trend prediction.This study applies the exponential smoothing method to process the initial data,calculates the relevant technical indicators as the characteristics to be selected,and proposes the D-RF-RS method to optimize random forest.As the random forest is an ensemble learning model and is closely related to decision tree,D-RF-RS method uses a decision tree to screen the importance of features,and obtains the effective strong feature set of the model as input.Then,the parameter combination of the model is optimized through random parameter search.The experimental results show that the average accuracy of random forest is increased by 0.17 after the above process optimization,which is 0.18 higher than the average accuracy of light gradient boosting machine model.Combined with the performance of the ROC curve and Precision–Recall curve,the stability of the model is also guaranteed,which further demonstrates the advantages of random forest in medium-and long-term trend prediction of the stock market.展开更多
The layered pavements usually exhibit complicated mechanical properties with the effect of complex material properties under external environment.In some cases,such as launching missiles or rockets,layered pavements a...The layered pavements usually exhibit complicated mechanical properties with the effect of complex material properties under external environment.In some cases,such as launching missiles or rockets,layered pavements are required to bear large impulse load.However,traditional methods cannot non-destructively and quickly detect the internal structural of pavements.Thus,accurate and fast prediction of the mechanical properties of layered pavements is of great importance and necessity.In recent years,machine learning has shown great superiority in solving nonlinear problems.In this work,we present a method of predicting the maximum deflection and damage factor of layered pavements under instantaneous large impact based on random forest regression with the deflection basin parameters obtained from falling weight deflection testing.The regression coefficient R^(2)of testing datasets are above 0.94 in the process of predicting the elastic moduli of structural layers and mechanical responses,which indicates that the prediction results have great consistency with finite element simulation results.This paper provides a novel method for fast and accurate prediction of pavement mechanical responses under instantaneous large impact load using partial structural parameters of pavements,and has application potential in non-destructive evaluation of pavement structure.展开更多
Power transformer is one of the most crucial devices in power grid.It is significant to determine incipient faults of power transformers fast and accurately.Input features play critical roles in fault diagnosis accura...Power transformer is one of the most crucial devices in power grid.It is significant to determine incipient faults of power transformers fast and accurately.Input features play critical roles in fault diagnosis accuracy.In order to further improve the fault diagnosis performance of power trans-formers,a random forest feature selection method coupled with optimized kernel extreme learning machine is presented in this study.Firstly,the random forest feature selection approach is adopted to rank 42 related input features derived from gas concentration,gas ratio and energy-weighted dissolved gas analysis.Afterwards,a kernel extreme learning machine tuned by the Aquila optimization algorithm is implemented to adjust crucial parameters and select the optimal feature subsets.The diagnosis accuracy is used to assess the fault diagnosis capability of concerned feature subsets.Finally,the optimal feature subsets are applied to establish fault diagnosis model.According to the experimental results based on two public datasets and comparison with 5 conventional approaches,it can be seen that the average accuracy of the pro-posed method is up to 94.5%,which is superior to that of other conventional approaches.Fault diagnosis performances verify that the optimum feature subset obtained by the presented method can dramatically improve power transformers fault diagnosis accuracy.展开更多
Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identificatio...Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identification of human body fluids,and has exhibited excellent performance in predicting single-source body fluids.The present study aims to develop a methylation SNaPshot multiplex system for body fluid identification,and accurately predict the mixture samples.In addition,the value of DNA methylation in the prediction of body fluid mixtures was further explored.Methods In the present study,420 samples of body fluid mixtures and 250 samples of single body fluids were tested using an optimized multiplex methylation system.Each kind of body fluid sample presented the specific methylation profiles of the 10 markers.Results Significant differences in methylation levels were observed between the mixtures and single body fluids.For all kinds of mixtures,the Spearman’s correlation analysis revealed a significantly strong correlation between the methylation levels and component proportions(1:20,1:10,1:5,1:1,5:1,10:1 and 20:1).Two random forest classification models were trained for the prediction of mixture types and the prediction of the mixture proportion of 2 components,based on the methylation levels of 10 markers.For the mixture prediction,Model-1 presented outstanding prediction accuracy,which reached up to 99.3%in 427 training samples,and had a remarkable accuracy of 100%in 243 independent test samples.For the mixture proportion prediction,Model-2 demonstrated an excellent accuracy of 98.8%in 252 training samples,and 98.2%in 168 independent test samples.The total prediction accuracy reached 99.3%for body fluid mixtures and 98.6%for the mixture proportions.Conclusion These results indicate the excellent capability and powerful value of the multiplex methylation system in the identification of forensic body fluid mixtures.展开更多
基金the National Natural Science Foundation of China(Grant 42177164)the Distinguished Youth Science Foundation of Hunan Province of China(2022JJ10073).
文摘As massive underground projects have become popular in dense urban cities,a problem has arisen:which model predicts the best for Tunnel Boring Machine(TBM)performance in these tunneling projects?However,performance level of TBMs in complex geological conditions is still a great challenge for practitioners and researchers.On the other hand,a reliable and accurate prediction of TBM performance is essential to planning an applicable tunnel construction schedule.The performance of TBM is very difficult to estimate due to various geotechnical and geological factors and machine specifications.The previously-proposed intelligent techniques in this field are mostly based on a single or base model with a low level of accuracy.Hence,this study aims to introduce a hybrid randomforest(RF)technique optimized by global harmony search with generalized oppositionbased learning(GOGHS)for forecasting TBM advance rate(AR).Optimizing the RF hyper-parameters in terms of,e.g.,tree number and maximum tree depth is the main objective of using the GOGHS-RF model.In the modelling of this study,a comprehensive databasewith themost influential parameters onTBMtogetherwithTBM AR were used as input and output variables,respectively.To examine the capability and power of the GOGHSRF model,three more hybrid models of particle swarm optimization-RF,genetic algorithm-RF and artificial bee colony-RF were also constructed to forecast TBM AR.Evaluation of the developed models was performed by calculating several performance indices,including determination coefficient(R2),root-mean-square-error(RMSE),and mean-absolute-percentage-error(MAPE).The results showed that theGOGHS-RF is a more accurate technique for estimatingTBMAR compared to the other applied models.The newly-developedGOGHS-RFmodel enjoyed R2=0.9937 and 0.9844,respectively,for train and test stages,which are higher than a pre-developed RF.Also,the importance of the input parameters was interpreted through the SHapley Additive exPlanations(SHAP)method,and it was found that thrust force per cutter is the most important variable on TBMAR.The GOGHS-RF model can be used in mechanized tunnel projects for predicting and checking performance.
基金financially supported by the National Natural Science Foundation of China(No.52174001)the National Natural Science Foundation of China(No.52004064)+1 种基金the Hainan Province Science and Technology Special Fund “Research on Real-time Intelligent Sensing Technology for Closed-loop Drilling of Oil and Gas Reservoirs in Deepwater Drilling”(ZDYF2023GXJS012)Heilongjiang Provincial Government and Daqing Oilfield's first batch of the scientific and technological key project “Research on the Construction Technology of Gulong Shale Oil Big Data Analysis System”(DQYT-2022-JS-750)。
文摘Real-time intelligent lithology identification while drilling is vital to realizing downhole closed-loop drilling. The complex and changeable geological environment in the drilling makes lithology identification face many challenges. This paper studies the problems of difficult feature information extraction,low precision of thin-layer identification and limited applicability of the model in intelligent lithologic identification. The author tries to improve the comprehensive performance of the lithology identification model from three aspects: data feature extraction, class balance, and model design. A new real-time intelligent lithology identification model of dynamic felling strategy weighted random forest algorithm(DFW-RF) is proposed. According to the feature selection results, gamma ray and 2 MHz phase resistivity are the logging while drilling(LWD) parameters that significantly influence lithology identification. The comprehensive performance of the DFW-RF lithology identification model has been verified in the application of 3 wells in different areas. By comparing the prediction results of five typical lithology identification algorithms, the DFW-RF model has a higher lithology identification accuracy rate and F1 score. This model improves the identification accuracy of thin-layer lithology and is effective and feasible in different geological environments. The DFW-RF model plays a truly efficient role in the realtime intelligent identification of lithologic information in closed-loop drilling and has greater applicability, which is worthy of being widely used in logging interpretation.
基金supported by the National Science Foundation of China(42107183).
文摘Driven piles are used in many geological environments as a practical and convenient structural component.Hence,the determination of the drivability of piles is actually of great importance in complex geotechnical applications.Conventional methods of predicting pile drivability often rely on simplified physicalmodels or empirical formulas,whichmay lack accuracy or applicability in complex geological conditions.Therefore,this study presents a practical machine learning approach,namely a Random Forest(RF)optimized by Bayesian Optimization(BO)and Particle Swarm Optimization(PSO),which not only enhances prediction accuracy but also better adapts to varying geological environments to predict the drivability parameters of piles(i.e.,maximumcompressive stress,maximum tensile stress,and blow per foot).In addition,support vector regression,extreme gradient boosting,k nearest neighbor,and decision tree are also used and applied for comparison purposes.In order to train and test these models,among the 4072 datasets collected with 17model inputs,3258 datasets were randomly selected for training,and the remaining 814 datasets were used for model testing.Lastly,the results of these models were compared and evaluated using two performance indices,i.e.,the root mean square error(RMSE)and the coefficient of determination(R2).The results indicate that the optimized RF model achieved lower RMSE than other prediction models in predicting the three parameters,specifically 0.044,0.438,and 0.146;and higher R^(2) values than other implemented techniques,specifically 0.966,0.884,and 0.977.In addition,the sensitivity and uncertainty of the optimized RF model were analyzed using Sobol sensitivity analysis and Monte Carlo(MC)simulation.It can be concluded that the optimized RF model could be used to predict the performance of the pile,and it may provide a useful reference for solving some problems under similar engineering conditions.
基金supported by the National Natural Science Foundation of China under Grant(Number:52105136)the Hong Kong Scholar program under Grant(Number:XJ2022013)China Postdoctoral Science Foundation under Grant(Number:2021M690290)Academic Excellence Foundation of BUAA under Grant(Number:BY2004103).
文摘Fatigue reliability-based design optimization of aeroengine structures involves multiple repeated calculations of reliability degree and large-scale calls of implicit high-nonlinearity limit state function,leading to the traditional direct Monte Claro and surrogate methods prone to unacceptable computing efficiency and accuracy.In this case,by fusing the random subspace strategy and weight allocation technology into bagging ensemble theory,a random forest(RF)model is presented to enhance the computing efficiency of reliability degree;moreover,by embedding the RF model into multilevel optimization model,an efficient RF-assisted fatigue reliability-based design optimization framework is developed.Regarding the low-cycle fatigue reliability-based design optimization of aeroengine turbine disc as a case,the effectiveness of the presented framework is validated.The reliabilitybased design optimization results exhibit that the proposed framework holds high computing accuracy and computing efficiency.The current efforts shed a light on the theory/method development of reliability-based design optimization of complex engineering structures.
文摘In the era of the Internet,widely used web applications have become the target of hacker attacks because they contain a large amount of personal information.Among these vulnerabilities,stealing private data through crosssite scripting(XSS)attacks is one of the most commonly used attacks by hackers.Currently,deep learning-based XSS attack detection methods have good application prospects;however,they suffer from problems such as being prone to overfitting,a high false alarm rate,and low accuracy.To address these issues,we propose a multi-stage feature extraction and fusion model for XSS detection based on Random Forest feature enhancement.The model utilizes RandomForests to capture the intrinsic structure and patterns of the data by extracting leaf node indices as features,which are subsequentlymergedwith the original data features to forma feature setwith richer information content.Further feature extraction is conducted through three parallel channels.Channel I utilizes parallel onedimensional convolutional layers(1Dconvolutional layers)with different convolutional kernel sizes to extract local features at different scales and performmulti-scale feature fusion;Channel II employsmaximum one-dimensional pooling layers(max 1D pooling layers)of various sizes to extract key features from the data;and Channel III extracts global information bi-directionally using a Bi-Directional Long-Short TermMemory Network(Bi-LSTM)and incorporates a multi-head attention mechanism to enhance global features.Finally,effective classification and prediction of XSS are performed by fusing the features of the three channels.To test the effectiveness of the model,we conduct experiments on six datasets.We achieve an accuracy of 100%on the UNSW-NB15 dataset and 99.99%on the CICIDS2017 dataset,which is higher than that of the existing models.
基金Under the auspices of National Natural Science Foundation of China(No.42071385)National Science and Technology Major Project of High Resolution Earth Observation System(No.79-Y50-G18-9001-22/23)。
文摘Automatically detecting Ulva prolifera(U.prolifera)in rainy and cloudy weather using remote sensing imagery has been a long-standing problem.Here,we address this challenge by combining high-resolution Synthetic Aperture Radar(SAR)imagery with the machine learning,and detect the U.prolifera of the South Yellow Sea of China(SYS)in 2021.The findings indicate that the Random Forest model can accurately and robustly detect U.prolifera,even in the presence of complex ocean backgrounds and speckle noise.Visual inspection confirmed that the method successfully identified the majority of pixels containing U.prolifera without misidentify-ing noise pixels or seawater pixels as U.prolifera.Additionally,the method demonstrated consistent performance across different im-ages,with an average Area Under Curve(AUC)of 0.930(+0.028).The analysis yielded an overall accuracy of over 96%,with an aver-age Kappa coefficient of 0.941(+0.038).Compared to the traditional thresholding method,Random Forest model has a lower estima-tion error of 14.81%.Practical application indicates that this method can be used in the detection of unprecedented U.prolifera in 2021 to derive continuous spatiotemporal changes.This study provides a potential new method to detect U.prolifera and enhances our under-standing of macroalgal outbreaks in the marine environment.
基金Under the auspices of National Natural Science Foundation of China(No.52079103)。
文摘Precise and timely prediction of crop yields is crucial for food security and the development of agricultural policies.However,crop yield is influenced by multiple factors within complex growth environments.Previous research has paid relatively little attention to the interference of environmental factors and drought on the growth of winter wheat.Therefore,there is an urgent need for more effective methods to explore the inherent relationship between these factors and crop yield,making precise yield prediction increasingly important.This study was based on four type of indicators including meteorological,crop growth status,environmental,and drought index,from October 2003 to June 2019 in Henan Province as the basic data for predicting winter wheat yield.Using the sparrow search al-gorithm combined with random forest(SSA-RF)under different input indicators,accuracy of winter wheat yield estimation was calcu-lated.The estimation accuracy of SSA-RF was compared with partial least squares regression(PLSR),extreme gradient boosting(XG-Boost),and random forest(RF)models.Finally,the determined optimal yield estimation method was used to predict winter wheat yield in three typical years.Following are the findings:1)the SSA-RF demonstrates superior performance in estimating winter wheat yield compared to other algorithms.The best yield estimation method is achieved by four types indicators’composition with SSA-RF)(R^(2)=0.805,RRMSE=9.9%.2)Crops growth status and environmental indicators play significant roles in wheat yield estimation,accounting for 46%and 22%of the yield importance among all indicators,respectively.3)Selecting indicators from October to April of the follow-ing year yielded the highest accuracy in winter wheat yield estimation,with an R^(2)of 0.826 and an RMSE of 9.0%.Yield estimates can be completed two months before the winter wheat harvest in June.4)The predicted performance will be slightly affected by severe drought.Compared with severe drought year(2011)(R^(2)=0.680)and normal year(2017)(R^(2)=0.790),the SSA-RF model has higher prediction accuracy for wet year(2018)(R^(2)=0.820).This study could provide an innovative approach for remote sensing estimation of winter wheat yield.yield.
基金support from the National Science and Technology Council of Taiwan(Contract Nos.111-2221 E-011081 and 111-2622-E-011019)the support from Intelligent Manufacturing Innovation Center(IMIC),National Taiwan University of Science and Technology(NTUST),Taipei,Taiwan,which is a Featured Areas Research Center in Higher Education Sprout Project of Ministry of Education(MOE),Taiwan(since 2023)was appreciatedWe also thank Wang Jhan Yang Charitable Trust Fund(Contract No.WJY 2020-HR-01)for its financial support.
文摘This study proposed a new real-time manufacturing process monitoring method to monitor and detect process shifts in manufacturing operations.Since real-time production process monitoring is critical in today’s smart manufacturing.The more robust the monitoring model,the more reliable a process is to be under control.In the past,many researchers have developed real-time monitoring methods to detect process shifts early.However,thesemethods have limitations in detecting process shifts as quickly as possible and handling various data volumes and varieties.In this paper,a robust monitoring model combining Gated Recurrent Unit(GRU)and Random Forest(RF)with Real-Time Contrast(RTC)called GRU-RF-RTC was proposed to detect process shifts rapidly.The effectiveness of the proposed GRU-RF-RTC model is first evaluated using multivariate normal and nonnormal distribution datasets.Then,to prove the applicability of the proposed model in a realmanufacturing setting,the model was evaluated using real-world normal and non-normal problems.The results demonstrate that the proposed GRU-RF-RTC outperforms other methods in detecting process shifts quickly with the lowest average out-of-control run length(ARL1)in all synthesis and real-world problems under normal and non-normal cases.The experiment results on real-world problems highlight the significance of the proposed GRU-RF-RTC model in modern manufacturing process monitoring applications.The result reveals that the proposed method improves the shift detection capability by 42.14%in normal and 43.64%in gamma distribution problems.
文摘目的:基于超高效液相色谱串联四极杆飞行时间质谱(UHPLC-QTOF-MS^(E))分析并经数字量化处理,结合随机森林(Random Forest,RF)算法构建数据辨识模型,以实现中华草龟、巴西龟、台湾龟、鳄鱼龟、鳖甲基原的数字化鉴定。方法:经样品预处理后,对不同来源、不同批次的龟甲进行UPLC-QTOF-MS^(E)分析,并以混合样品为基准进行峰位校正、提取并经量化处理,获取反映多肽离子信息的精确质量数-保留时间数据对(Exact Mass Retention Time,EMRT)。然后基于信息增益率的特征筛选获取重要多肽离子信息,结合随机森林(RF)算法进行数据建模,同时基于内部交叉验证中的准确率(Acc)、精确率(P)、曲线下面积(AUC)等参数进行模型评价。最后基于最优模型进行龟甲基原的鉴定验证分析。结果:基于信息增益率的特征筛选,得到71个特征多肽信息,建立的RF模型具有优秀的辨识效果,准确率、精确率以及AUC均大于0.950且外部鉴定验证的正确率为100.0%。结论:基于UHPLC-QTOF-MS^(E)分析,并结合RF算法能够高效准确地实现不同来源龟甲基原的数字化鉴定,可为龟甲的质量控制及基原考证提供参考和帮助。
基金supported by the Guangdong Major Project of Basic and Applied Basic Research under Grant No. 2021B0301030006the computational resources from SYSU and the National Supercomputer Center in Guangzhou。
文摘Information on the decay process of nuclides in the superheavy region is critical in investigating new elements beyond oganesson and the island of stability.This paper presents the application of a random forest algorithm to examine the competition among different decay modes in the superheavy region,includingα decay,β^(-)decay,β^(+)decay,electron capture and spontaneous fission.The observed half-lives and dominant decay mode are well reproduced.The dominant decay mode of 96.9%of the nuclei beyond ^(212) Po is correctly obtained.Further,α decay is predicted to be the dominant decay mode for isotopes in new elements Z=119-122,except for spontaneous fission in certain even–even elements owing to the increased Coulomb repulsion and odd–even effect.The predicted half-lives demonstrate the existence of a long-lived spontaneous fission island southwest of ^(298) Fl caused by the competition between the fission barrier and Coulomb repulsion.A better understanding of spontaneous fission,particularly beyond ^(286)Fl,is crucial in the search for new elements and the island of stability.
基金supported by the National Natural Science Foundation of China(61873006)Beijing Natural Science Foundation(4204087,4212040).
文摘Acid production with flue gas is a complex nonlinear process with multiple variables and strong coupling.The operation data is an important basis for state monitoring,optimal control,and fault diagnosis.However,the operating environment of acid production with flue gas is complex and there is much equipment.The data obtained by the detection equipment is seriously polluted and prone to abnormal phenomena such as data loss and outliers.Therefore,to solve the problem of abnormal data in the process of acid production with flue gas,a data cleaning method based on improved random forest is proposed.Firstly,an outlier data recognition model based on isolation forest is designed to identify and eliminate the outliers in the dataset.Secondly,an improved random forest regression model is established.Genetic algorithm is used to optimize the hyperparameters of the random forest regression model.Then the optimal parameter combination is found in the search space and the trend of data is predicted.Finally,the improved random forest data cleaning method is used to compensate for the missing data after eliminating abnormal data and the data cleaning is realized.Results show that the proposed method can accurately eliminate and compensate for the abnormal data in the process of acid production with flue gas.The method improves the accuracy of compensation for missing data.With the data after cleaning,a more accurate model can be established,which is significant to the subsequent temperature control.The conversion rate of SO_(2) can be further improved,thereby improving the yield of sulfuric acid and economic benefits.
基金supported by National Natural Science Foundation of China(No.62101298)Collaborative Education Project between Industry and Academia,China(22050609312501)。
文摘60 GHz millimeter wave(mmWave)system provides extremely high time resolution and multipath components(MPC)separation and has great potential to achieve high precision in the indoor positioning.However,the ranging data is often contaminated by non-line-of-sight(NLOS)transmission.First,six features of 60GHz mm Wave signal under LOS and NLOS conditions are evaluated.Next,a classifier constructed by random forest(RF)algorithm is used to identify line-of-sight(LOS)or NLOS channel.The identification mechanism has excellent generalization performance and the classification accuracy is over 97%.Finally,based on the identification results,a residual weighted least squares positioning method is proposed.All ranging information including that under NLOS channels is fully utilized,positioning failure caused by insufficient LOS links can be avoided.Compared with the conventional least squares approach,the positioning error of the proposed algorithm is reduced by 49%.
文摘AIM:To analyze ultrasound biomicroscopy(UBM)images using random forest network to find new features to make predictions about vault after implantable collamer lens(ICL)implantation.METHODS:A total of 450 UBM images were collected from the Lixiang Eye Hospital to provide the patient’s preoperative parameters as well as the vault of the ICL after implantation.The vault was set as the prediction target,and the input elements were mainly ciliary sulcus shape parameters,which included 6 angular parameters,2 area parameters,and 2 parameters,distance between ciliary sulci,and anterior chamber height.A random forest regression model was applied to predict the vault,with the number of base estimators(n_estimators)of 2000,the maximum tree depth(max_depth)of 17,the number of tree features(max_features)of Auto,and the random state(random_state)of 40.0.RESULTS:Among the parameters selected in this study,the distance between ciliary sulci had a greater importance proportion,reaching 52%before parameter optimization is performed,and other features had less influence,with an importance proportion of about 5%.The importance of the distance between the ciliary sulci increased to 53% after parameter optimization,and the importance of angle 3 and area 1 increased to 5% and 8%respectively,while the importance of the other parameters remained unchanged,and the distance between the ciliary sulci was considered the most important feature.Other features,although they accounted for a relatively small proportion,also had an impact on the vault prediction.After parameter optimization,the best prediction results were obtained,with a predicted mean value of 763.688μm and an actual mean value of 776.9304μm.The R²was 0.4456 and the root mean square error was 201.5166.CONCLUSION:A study based on UBM images using random forest network can be performed for prediction of the vault after ICL implantation and can provide some reference for ICL size selection.
基金supported by the Elite Scholar Program of Northwest A&F University (Grant No.Z111022001)the Research Fund of Department of Transport of Shannxi Province (Grant No.22-23K)the Student Innovation and Entrepreneurship Training Program of China (Project Nos.S202110712555 and S202110712534).
文摘A huge number of old arch bridges located in rural regions are at the peak of maintenance.The health monitoring technology of the long-span bridge is hardly applicable to the small-span bridge,owing to the absence of technical resources and sufficient funds in rural regions.There is an urgent need for an economical,fast,and accurate damage identification solution.The authors proposed a damage identification system of an old arch bridge implemented with amachine learning algorithm,which took the vehicle-induced response as the excitation.A damage index was defined based on wavelet packet theory,and a machine learning sample database collecting the denoised response was constructed.Through comparing three machine learning algorithms:Back-Propagation Neural Network(BPNN),Support Vector Machine(SVM),and Random Forest(R.F.),the R.F.damage identification model were found to have a better recognition ability.Finally,the Particle Swarm Optimization(PSO)algorithm was used to optimize the number of subtrees and split features of the R.F.model.The PSO optimized R.F.model was capable of the identification of different damage levels of old arch bridges with sensitive damage index.The proposed framework is practical and promising for the old bridge’s structural damage identification in rural regions.
基金National Natural Science Foundation of China,Grant/Award Numbers:61673084,National Natural Science Foundation of ChinaThe Fundamental Research Foundation for Universities of Heilongjiang Province,Grant/Award Number:LGYC2018JC017。
文摘As a complex hot problem in the financial field,stock trend forecasting uses a large amount of data and many related indicators;hence it is difficult to obtain sustainable and effective results only by relying on empirical analysis.Researchers in the field of machine learning have proved that random forest can form better judgements on this kind of problem,and it has an auxiliary role in the prediction of stock trend.This study uses historical trading data of four listed companies in the USA stock market,and the purpose of this study is to improve the performance of random forest model in medium-and long-term stock trend prediction.This study applies the exponential smoothing method to process the initial data,calculates the relevant technical indicators as the characteristics to be selected,and proposes the D-RF-RS method to optimize random forest.As the random forest is an ensemble learning model and is closely related to decision tree,D-RF-RS method uses a decision tree to screen the importance of features,and obtains the effective strong feature set of the model as input.Then,the parameter combination of the model is optimized through random parameter search.The experimental results show that the average accuracy of random forest is increased by 0.17 after the above process optimization,which is 0.18 higher than the average accuracy of light gradient boosting machine model.Combined with the performance of the ROC curve and Precision–Recall curve,the stability of the model is also guaranteed,which further demonstrates the advantages of random forest in medium-and long-term trend prediction of the stock market.
基金Project supported in part by the National Natural Science Foundation of China(Grant No.12075168)the Fund from the Science and Technology Commission of Shanghai Municipality(Grant No.21JC1405600)。
文摘The layered pavements usually exhibit complicated mechanical properties with the effect of complex material properties under external environment.In some cases,such as launching missiles or rockets,layered pavements are required to bear large impulse load.However,traditional methods cannot non-destructively and quickly detect the internal structural of pavements.Thus,accurate and fast prediction of the mechanical properties of layered pavements is of great importance and necessity.In recent years,machine learning has shown great superiority in solving nonlinear problems.In this work,we present a method of predicting the maximum deflection and damage factor of layered pavements under instantaneous large impact based on random forest regression with the deflection basin parameters obtained from falling weight deflection testing.The regression coefficient R^(2)of testing datasets are above 0.94 in the process of predicting the elastic moduli of structural layers and mechanical responses,which indicates that the prediction results have great consistency with finite element simulation results.This paper provides a novel method for fast and accurate prediction of pavement mechanical responses under instantaneous large impact load using partial structural parameters of pavements,and has application potential in non-destructive evaluation of pavement structure.
基金support of national natural science foundation of China(No.52067021)natural science foundation of Xinjiang(2022D01C35)+1 种基金excellent youth scientific and technological talents plan of Xinjiang(No.2019Q012)major science and technology special project of Xinjiang Uygur Autonomous Region(2022A01002-2).
文摘Power transformer is one of the most crucial devices in power grid.It is significant to determine incipient faults of power transformers fast and accurately.Input features play critical roles in fault diagnosis accuracy.In order to further improve the fault diagnosis performance of power trans-formers,a random forest feature selection method coupled with optimized kernel extreme learning machine is presented in this study.Firstly,the random forest feature selection approach is adopted to rank 42 related input features derived from gas concentration,gas ratio and energy-weighted dissolved gas analysis.Afterwards,a kernel extreme learning machine tuned by the Aquila optimization algorithm is implemented to adjust crucial parameters and select the optimal feature subsets.The diagnosis accuracy is used to assess the fault diagnosis capability of concerned feature subsets.Finally,the optimal feature subsets are applied to establish fault diagnosis model.According to the experimental results based on two public datasets and comparison with 5 conventional approaches,it can be seen that the average accuracy of the pro-posed method is up to 94.5%,which is superior to that of other conventional approaches.Fault diagnosis performances verify that the optimum feature subset obtained by the presented method can dramatically improve power transformers fault diagnosis accuracy.
基金supported by the grants from the Natural Science Foundation of Hubei Province(No.2020CFB780)the Fundamental Research Funds for the Central Universities(No.2017KFYXJJ020).
文摘Objective Body fluid mixtures are complex biological samples that frequently occur in crime scenes,and can provide important clues for criminal case analysis.DNA methylation assay has been applied in the identification of human body fluids,and has exhibited excellent performance in predicting single-source body fluids.The present study aims to develop a methylation SNaPshot multiplex system for body fluid identification,and accurately predict the mixture samples.In addition,the value of DNA methylation in the prediction of body fluid mixtures was further explored.Methods In the present study,420 samples of body fluid mixtures and 250 samples of single body fluids were tested using an optimized multiplex methylation system.Each kind of body fluid sample presented the specific methylation profiles of the 10 markers.Results Significant differences in methylation levels were observed between the mixtures and single body fluids.For all kinds of mixtures,the Spearman’s correlation analysis revealed a significantly strong correlation between the methylation levels and component proportions(1:20,1:10,1:5,1:1,5:1,10:1 and 20:1).Two random forest classification models were trained for the prediction of mixture types and the prediction of the mixture proportion of 2 components,based on the methylation levels of 10 markers.For the mixture prediction,Model-1 presented outstanding prediction accuracy,which reached up to 99.3%in 427 training samples,and had a remarkable accuracy of 100%in 243 independent test samples.For the mixture proportion prediction,Model-2 demonstrated an excellent accuracy of 98.8%in 252 training samples,and 98.2%in 168 independent test samples.The total prediction accuracy reached 99.3%for body fluid mixtures and 98.6%for the mixture proportions.Conclusion These results indicate the excellent capability and powerful value of the multiplex methylation system in the identification of forensic body fluid mixtures.