-- SandraCastro - 2015-04-15

Casa HPC/Parallelization meeting agenda/minutes


Thursday [16/04/2015], [ESO Centaurus, C.2.01], [15:00 UT]


How to connect

  • Dial in: +49 89 6834
  • Video connection: Use this one (46103@134.171.42.27).

Attendees:

ESO: Sandra, Julian, Justo

Socorro: Martin, Jeff, Lindsey, Kumar, Tak, James

CV: Akeem, Mark, Andy

Agenda

  1. Open items from last weeks:
    • Find a solution for having the CASA trunk build on the cluster without the use of tricks and patches as it is currently done.
    • Mark to clarify what are the differences between casa and casapy scripts and their use-cases.
  2. Where are we on the tile caching solution for MMS cases?
  3. Setjy failures in MMS, not in MS when running the EVLA pipeline. One logfile is in: /lustre/scastro/testpipe/evla/parallel/14B_14subms_casapy-20150410-145444.log
    2015-04-11 04:37:23   SEVERE   setjy::::@nmpost007::MPIServer-10   Use standard="manual" to set the specified fluxdensity.
    2015-04-11 04:37:23   SEVERE   setjy::::@nmpost007::MPIServer-10   An error occurred running task setjy: Use standard="manual" to set the specified fluxdensity.
  4. Justo has updated the list of system requirements for the HPC development. Please have a look at https://safe.nrao.edu/wiki/bin/view/HPC/SystemRequirements and send comments.
  5. I have started to update the section use cases applicable to starting and running a parallelized version of CASA on various different clusters, on the user's point of view. https://safe.nrao.edu/wiki/bin/view/HPC/ConceptOfOperation
  6. Round the table to get the status of what each one is doing.
  7. Do we want to have another HPC lecture soon? On which topic?

Minutes

  • It seems that the degradation in performance of gaincal on MMS seems to happen in lustre, but not in local server. James doubts it is related to caching issues in lustre. James will check in lustre, comparing file systems. Julian thinks itÂ’s the same problem we've seen since November in which the task makes small reads many times. The tile caching fix from George has not fixed anything in the MMS cases. Jeff suggests to multiply the cache by 10 and verify if it affects the run in the same proportion. Julian suggests to look also for fsync calls in gaincal, which was a problem previously in importasdm.
  • Tak will verify the failures in setjy when running the EVLA pipeline on MMS.
  • Next steps for everyone:
    • Jeff will continue the discussions on the system requirements with ESO team until we reach an agreement. Stewart has been put in the loop too.
    • James will look at the gaincal issue in lustre.
    • Tak will look at setjy failure and cube parallelization.
    • Lindsey will look at pipeline performance issues and plotms efficiency.
    • Martin had a meeting with Urvashi and Jeff and has a conceptual model in his mind of what to do for the parallelization of IC in tclean.
    • Justo will look at using strace and i/o profilers to check the problems in gaincal. Will also iterate with Jeff on the system requirements on tiers 2,3.
    • Julian will change the cache values in gaincal and check the results. He has also to check why the MultiFile is slower than standard implementation. He suggested that the use of MultiFile should added to partition at the time of the MMS creation, with some heuristics to determine when to use it or not.
    • Sandra will continue testing the ALMA pipeline: fix from Pam on plotms write-lock; lazy import + partition. She will also udpate the MMS document with the information about the balanced mode in partition.

Action Item List

Topic revision: r4 - 2015-04-20, SandraCastro
This site is powered by FoswikiCopyright © by the contributing authors. All material on this collaboration platform is the property of the contributing authors.
Ideas, requests, problems regarding NRAO Public Wiki? Send feedback