AQX: Explaining Air Quality Forecast for Verifying Domain Knowledge using Feature Importance Visualization
Authors
Title of the Paper
AQX: Explaining Air Quality Forecast for Verifying Domain Knowledge using Feature Importance Visualization
Paper Information
- Research Area: Explainable Artificial Intelligence (XAI), Visual Analytics, Air Quality Prediction
- Keywords: Explainable Artificial Intelligence (XAI), Machine Learning (ML), Air Pollution, Air Quality Prediction, Spatio-Temporal Data, Feature Importance, Visual Analytics, Domain Knowledge Verification
Research Background and Problem
-
What problems or challenges did the authors identify?
With industrialization, air pollution has severely impacted human health. Although machine learning models (ML models) have the potential to make high-precision predictions using historical data, their "black-box" nature has raised skepticism among environmental science experts. This lack of transparency limits the application of ML models in practical policymaking. In contrast, traditional CMAQ simulation models, despite sometimes being less accurate, are widely accepted due to their transparent rule-based operation. -
Why is this problem important?
The direct impact of air pollution on human health makes air quality prediction a critical area of scientific research. Explaining machine learning models and verifying whether their prediction processes align with existing domain knowledge can enhance trust in these models and promote their adoption in practical applications. -
Research Motivation and Related Work
The authors argue that verifying whether machine learning models' learning processes align with domain knowledge can address experts' skepticism. While the XAI field has proposed several methods to tackle the transparency issues of machine learning models, there is limited research on how to use XAI tools to verify domain knowledge. Therefore, this paper focuses on providing a visual analytics system to support domain knowledge verification.
Solution
-
What methods or solutions did the authors propose?
The authors developed a visual analytics system named AQX to verify whether machine learning models' learning processes for air pollution prediction align with domain knowledge. AQX uses feature importance visualization to demonstrate the contribution of model inputs to prediction results, while also incorporating model performance comparisons and original wind data to validate domain knowledge. -
What are the innovative aspects of this solution?
- Introducing a feature importance visualization tool to display the contributions of input features across different temporal and spatial dimensions.
- Comparing traditional simulation models (CMAQ) with machine learning models to identify areas where model learning aligns or diverges from real-world scenarios.
- Utilizing animations to display original wind trajectory data, providing additional support for validating model learning.
-
What are the implementation steps and key technologies used?
- Building the Prediction Model: Developing a machine learning model based on the ConvLSTM architecture to predict air pollution concentrations.
- Feature Importance Calculation: Extracting the contributions of different inputs to model predictions using gradient calculation methods.
- Data Visualization Design: Designing a system interface with five views (overview, performance view, raw data view, feature contribution view, spatio-temporal contribution view).
- User Interaction Support: Providing selection, filtering, and other functionalities to facilitate in-depth exploration by experts.
Research Outcomes
-
What specific results were achieved?
AQX helped domain experts verify whether machine learning models' learning processes align with existing knowledge and analyze model performance in special cases such as extreme weather conditions. Additionally, the system demonstrated its effectiveness and applicability through two case studies and expert interviews. -
What advantages does it have compared to existing solutions?
- Multi-level feature importance visualization (temporal and spatial dimensions).
- Comparison of performance and spatial consistency between traditional models and machine learning models.
- Original wind trajectory animations supporting the domain experts' verification process.
-
What were the experimental or evaluation results?
Two domain experts evaluated AQX and verified core domain knowledge, such as the impact of wind on air quality, correlations between air pollutants, and dependencies on spatio-temporal characteristics. These results helped experts better understand the performance and rationality of ML models. -
Limitations and Future Directions
- Scalability Issues: The visualization design may face challenges when handling a large number of input features.
- Cognitive Load: Users need to retain a significant amount of information to support cross-month or cross-dimensional comparative analysis.
- Involvement of More Domain Experts: The number of experts involved in the study and evaluation is currently limited, and future work could expand to a broader user base.
- Application to Other Domains: AQX could be extended to areas such as precipitation forecasting and aerodynamics, requiring adjustments to meet specific domain needs.
By verifying domain knowledge, AQX provides an essential tool for enhancing the transparency and trustworthiness of machine learning models, while also amplifying the potential impact of explainable AI in environmental science. Future research could further explore how expert feedback can improve system usability and implement similar knowledge verification functionalities in other complex domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Are feature importances from machine learning predictions consistent with existing knowledge in air pollution?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
- How can visualization tools verify the alignment between machine learning models and domain knowledge?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
- How can spatiotemporal data improve the transparency of air quality prediction models?Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
Practical Problems
1- Environmental science experts do not trust the black-box predictions of machine learning models.Category: Transparency, Auditability, and Trust Calibration MechanismsSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)