Artificial Intelligence Value Alignment via Inverse Reinforcement Learning
Carregando...
Arquivos
Data
2023-12-19
Autores
Orientador(res)
Mesquita, Diego Parente Paiva
Métricas
Título da Revista
ISSN da Revista
Título de Volume
Resumo
This dissertation presents a comprehensive exploration of the field of artificial intelligence (AI) alignment, emphasizing the integration of human values and ethics into AI systems. The research synthesizes a broad range of academic literature to elucidate the principles, methodologies, challenges, and ethical implications inherent in AI alignment. Central to this discussion are the principles of robustness, interpretability, controllability, and ethicality, which are critical for the development of AI systems that are not only technically proficient but also ethically aligned with human values. The dissertation delves into the dual aspects of AI alignment: forward alignment, which focuses on embedding human values during the AI training phase, and backward alignment, emphasizing ongoing governance and verification post-deployment. A key challenge identified is the integration of complex and often subjective human values into computational models, highlighting limitations in current methodologies like Inverse Reinforcement Learning (IRL). The ethical and safety considerations in IRL are critically examined, underscoring the need for a balance between technological advancement and ethical integrity. The dissertation advocates for methodological advancements, hybrid approaches combining empirical data and ethical reasoning, addressing data biases, and establishing robust governance frameworks. Future research directions identified include methodological innovations, addressing data biases, and the need for interdisciplinary collaborations to tackle the multifaceted challenges of AI alignment. This research concludes that AI alignment is vital for addressing existential risks posed by AI, as it ensures AI development with human values and ethics, a key step in preventing AI technologies from diverging in potentially harmful ways. By emphasizing the principles of ethical alignment in AI systems, it contributes to mitigating the risks of unaligned powerful AIs and ensuring a safe, harmonious coexistence between AI and humanity.
