Hi all,

SInce last Wednesday (2016-12-07) we at jao started having reports of some jaopost client machines (around 10 in 4 days) that got frozen during Pipeline Reductions over Lustre.

The machines doesn't respond to ping or ssh, and when we check them through iDrac or connecting a Screen we got following message:

LustreError: 22231:0:(mdc_locks.c:918:mdc_enqueue()) Skipped 1 previous similar message

casaplotms[33429]: segfault at 8 ip 00007f3fcfb3bd00 sp 00007fffa54cd748 error 4 in libmsvis.so.1364.178.125[7f3fcf76d000+4df000]

LustreError: 33373:0 (mdc_locks.c:918:mdc_enqueue()) ldlm_cli_enqueue: -2

LustreError: 33373:0 (mdc_locks.c:918:mdc_enqueue()) Skipped 1 previous similar message

Kernel panic - not syncing: Attempted to kill init!

Call Trace:

[<ffffffff815284fc>] ? panic+0xa7/0x16f

[<ffffffff813321a6>] ? get_current_tty+0x66/0x70

[<ffffffff81077332>] ? do_exit+0x862/0x870

[<ffffffff81088f5d>] ? __sigqueue_free+0x3d/0x50

[<ffffffff81077398>] ? do_group_exit+0x58/0xd0

[<ffffffff8108cd46>] ? get_signal_to_deliver+0x1f6/0x460

[<ffffffff8100a256>] ? do_signal+0x75/0x800

[<ffffffff8109b39c>] ? remove_wait_queue+0x3c/0x50

[<ffffffff81010000>] ? dump_trace+0x80/0x3b0

[<ffffffff81049e17>] ? is_prefetch=0x1a7/0x230

[<ffffffff81076408>] ? sys_waitid+0xa8/0x1f0

[<ffffffff8100aa80>] ? do_notify_resume+0x90/0xc0

[<ffffffff8100badc>] ? retint_signal+0x48/0x8c

-- BernardoMalet - 2016-12-12
Topic revision: r1 - 2016-12-12, BernardoMalet
This site is powered by FoswikiCopyright © by the contributing authors. All material on this collaboration platform is the property of the contributing authors.
Ideas, requests, problems regarding NRAO Public Wiki? Send feedback