Logically consistent adversarial attacks for soft theorem provers
File(s)
Author(s)
Gaskell, Alexander
Miao, Yishu
Toni, Francesca
Specia, Lucia
Type
Conference Paper
Abstract
Recent efforts within the AI community have
yielded impressive results towards “soft theorem
proving” over natural language sentences using lan-
guage models. We propose a novel, generative
adversarial framework for probing and improving
these models’ reasoning capabilities. Adversarial
attacks in this domain suffer from the logical in-
consistency problem, whereby perturbations to the
input may alter the label. Our Logically consis-
tent AdVersarial Attacker, LAVA, addresses this by
combining a structured generative process with a
symbolic solver, guaranteeing logical consistency.
Our framework successfully generates adversarial
attacks and identifies global weaknesses common
across multiple target models. Our analyses reveal
naive heuristics and vulnerabilities in these mod-
els’ reasoning capabilities, exposing an incomplete
grasp of logical deduction under logic programs.
Finally, in addition to effective probing of these
models, we show that training on the generated
samples improves the target model’s performance.
yielded impressive results towards “soft theorem
proving” over natural language sentences using lan-
guage models. We propose a novel, generative
adversarial framework for probing and improving
these models’ reasoning capabilities. Adversarial
attacks in this domain suffer from the logical in-
consistency problem, whereby perturbations to the
input may alter the label. Our Logically consis-
tent AdVersarial Attacker, LAVA, addresses this by
combining a structured generative process with a
symbolic solver, guaranteeing logical consistency.
Our framework successfully generates adversarial
attacks and identifies global weaknesses common
across multiple target models. Our analyses reveal
naive heuristics and vulnerabilities in these mod-
els’ reasoning capabilities, exposing an incomplete
grasp of logical deduction under logic programs.
Finally, in addition to effective probing of these
models, we show that training on the generated
samples improves the target model’s performance.
Date Acceptance
2022-04-20
Citation
Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, pp.4129-4135
ISBN
978-1-956792-00-3
Publisher
International Joint Conferences on Artificial Intelligence
Start Page
4129
End Page
4135
Journal / Book Title
Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence
Copyright Statement
Copyright © 2022 International Joint Conferences on Artificial Intelligence. All rights reserved. No part of this book may be reproduced in any form by any electronic or mechanical means (including photocopying, recording, or information storage and retrieval) without permission in writing from the publisher.
Identifier
https://www.ijcai.org/proceedings/2022/573
Source
31st International Joint Conference on Artificial Intelligence and the 25th European Conference on Artificial Intelligence
Publication Status
Published
Start Date
2022-07-23
Finish Date
2022-07-29
Coverage Spatial
Vienna, Austria
Date Publish Online
2022-07-23
