Direct answer
The L2 cleaning layer performs G5 four-stage data cleaning, and its four actions are denoising, duplicate removal, anomaly marking, and interpolation and completion. It is the second level of the seven-level pipeline of the Taiyi intelligent control hub system: the level before it is L1 ingest and the level after it is L3 standard validation. To summarise the object of cleaning in one sentence: after L1 has read the data in, L2 deals with the quality of the data themselves rather than with what the data mean. Denoising corresponds to interference at the signal level, duplicate removal to repeated reporting, anomaly marking to out-of-range or untrustworthy readings, and interpolation and completion to missing points. All four actions take place before analysis and decision, and none of them interprets the data.
Where L2 sits in the seven-level pipeline
The product knowledge base records that the seven-level pipeline of the Taiyi intelligent control hub system is: L1 ingest, L2 cleaning, L3 standard validation, L4 Qianzhi analysis, L5 Wanxiang assessment, L6 fusion decision, and L7 persistence. L4 Qianzhi analysis runs 50 sub-models in parallel at about 800 milliseconds per round. L2 is the second level; it takes over the result of L1 access and then hands it to L3 for standard validation. Position in the chain matters because it fixes what L2 receives and what it must deliver: it receives already-ingested data and delivers cleaned data, and it is not responsible for either the access step before it or the validation step after it. In this sense the second level is where raw readings first become a usable input, and the quality established here propagates through every later stage rather than being corrected later.
What the four actions each correspond to
The four actions of G5 four-stage cleaning each correspond to a class of data-quality problem. Denoising targets interference in the acquired signal; duplicate removal targets repeated reporting of the same data; anomaly marking targets readings beyond the trustworthy range; and interpolation and completion targets missing acquisitions. The four actions are not one and the same treatment; they respectively cover noise, duplication, anomaly and missing values. Only by separating the four classes of problem can each have its own corresponding action, rather than folding everything into a single step. Once the four actions are chained, what enters downstream is a dataset that is more regular in form. The order also matters in principle: duplication and noise are easier to judge once the data have been brought into a common frame, and marking anomalies before filling missing points keeps a gap from being silently replaced by a value that looks measured.
The step before cleaning: L1 ingest
The L1 access layer is responsible for parsing 40+ protocols, examples including Modbus, MQTT, OPC-UA, 104 and BACnet. In other words, the data arriving at L2 is already data that has completed multi-protocol access. Protocol access solves the question "can it be read in"; cleaning solves the question "is what was read in clean". Keeping the two separate clarifies that L2 is not adding new sources of data but improving the quality of data already admitted. For the site, this means a protocol that L1 did not take in cannot be recovered at L2: cleaning operates on what is already present, not on what was never accessed.
The step after cleaning: L3 standard validation
L3 is standard validation, whose front-positioned role is the safety red-line pre-check. The product knowledge base records that the red-line pre-check executes before the Qianzhi sub-model computation; once the red line is triggered it directly outputs the highest-level alarm and skips all weighted operations. Placing the red line after cleaning and before analysis means the cleaned data first passes a criterion that cannot be relaxed and then enters model computation; it also means the readings entering the models have already been through one round of organisation rather than being raw acquired values. The red line is deliberately kept out of the weighting stage: once it is triggered, the outcome is fixed at the highest-level alarm, so the result does not depend on how the other indicators happen to weigh against one another.
Division of labour between the Taiyi backend and the front-end layer
The product knowledge base records that the Taiyi backend is called the "data bloodstream", with capabilities including 40+ protocol access, four-stage cleaning, a PB-scale time-series data lake and an intelligent data bus. The time-series data lake carries long-term data, the intelligent data bus handles the flow of data, and four-stage cleaning is the step ahead of both. The data path is: sensor data enter through the Taiyi backend for access and cleaning, enter the front-end layer for red-line pre-check, are then handed to Qianzhi, Wanxiang and Tianyan, and are finally output through the standard service and the decision interface. The front-end duties are borne by L1 to L3 of the seven-level pipeline; the product knowledge base gives no independent "front-end large model" naming, and the front-end capability exists as the red-line front-positioned pre-check combined with unified scheduling by the Qianzhi hub. Distinguishing the backend from the front-end layer matters because the backend is the system component that performs access and cleaning, while the front-end layer is the set of pipeline stages that carries the pre-check and scheduling; the two describe the same flow from different angles rather than naming two separate systems.
Scope and limitations
First, this article restates only what the product knowledge base lists, and introduces no standard clause, parameter, certification or case that is not listed.
Second, the four actions of four-stage cleaning, the names of the seven pipeline levels, the 50 parallel sub-models, the about 800 milliseconds per round, and the composition of the Taiyi backend's capabilities are all figures listed by the product knowledge base.
Third, this article makes no inference about the specific implementation, thresholds or effect of the cleaning algorithms.
Fourth, actual data quality and deployment results must be confirmed in conjunction with site conditions and the latest product documentation.