<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">OS</journal-id><journal-title-group>
    <journal-title>Ocean Science</journal-title>
    <abbrev-journal-title abbrev-type="publisher">OS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Ocean Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1812-0792</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/os-16-831-2020</article-id><title-group><article-title>An approach to the verification of high-resolution ocean models using
spatial methods</article-title><alt-title>Verification of high-resolution ocean models using
spatial methods</alt-title>
      </title-group><?xmltex \runningtitle{Verification of high-resolution ocean models using
spatial methods}?><?xmltex \runningauthor{R.~Crocker et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Crocker</surname><given-names>Ric</given-names></name>
          <email>ric.crocker@metoffice.gov.uk</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Maksymczuk</surname><given-names>Jan</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Mittermaier</surname><given-names>Marion</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-4752-3135</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Tonani</surname><given-names>Marina</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-5452-3251</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Pequignet</surname><given-names>Christine</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-3057-8300</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Verification, Impacts and Post-Processing, Weather Science, Met
Office, Exeter, EX1 3PB, UK</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Ocean Forecasting Research &amp; Development, Weather Science, Met
Office, Exeter, EX1 3PB, UK</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Ric Crocker (ric.crocker@metoffice.gov.uk)</corresp></author-notes><pub-date><day>17</day><month>July</month><year>2020</year></pub-date>
      
      <volume>16</volume>
      <issue>4</issue>
      <fpage>831</fpage><lpage>845</lpage>
      <history>
        <date date-type="received"><day>13</day><month>February</month><year>2020</year></date>
           <date date-type="rev-request"><day>28</day><month>February</month><year>2020</year></date>
           <date date-type="rev-recd"><day>28</day><month>May</month><year>2020</year></date>
           <date date-type="accepted"><day>16</day><month>June</month><year>2020</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2020 Ric Crocker et al.</copyright-statement>
        <copyright-year>2020</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020.html">This article is available from https://os.copernicus.org/articles/16/831/2020/os-16-831-2020.html</self-uri><self-uri xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020.pdf">The full text article is available as a PDF file from https://os.copernicus.org/articles/16/831/2020/os-16-831-2020.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e123">The Met Office currently runs two operational ocean forecasting
configurations for the  North West European Shelf: an eddy-permitting model
with a resolution of 7 km (AMM7) and an eddy-resolving model at 1.5 km
(AMM15).</p>
    <p id="d1e126">Whilst qualitative assessments have demonstrated the benefits brought by the
increased resolution of AMM15, particularly in the ability to resolve
finer-scale features, it has been difficult to show this quantitatively,
especially in forecast mode. Applications of typical assessment metrics such
as the root mean square error have been inconclusive, as the high-resolution
model tends to be penalised more severely, referred to as the double-penalty
effect. This effect occurs in point-to-point comparisons whereby features
correctly forecast but misplaced with respect to the observations are
penalised twice: once for not occurring at the observed location, and
secondly for occurring at the forecast location, where they have not been
observed.</p>
    <p id="d1e129">An exploratory assessment of sea surface temperature (SST) has been made at
in situ observation locations using a
single-observation neighbourhood-forecast (SO-NF) spatial verification
method known as the High-Resolution Assessment (HiRA) framework. The primary
focus of the assessment was to capture important aspects of methodology to
consider when applying the HiRA framework. Forecast grid points within
neighbourhoods centred on the observing location are considered as pseudo
ensemble members, so that typical ensemble and probabilistic forecast
verification metrics such as the continuous ranked probability score (CRPS)
can be utilised. It is found that through the application of HiRA it is
possible to identify improvements in the higher-resolution model which were
not apparent using typical grid-scale assessments.</p>
    <p id="d1e132">This work suggests that future comparative assessments of ocean models with
different resolutions would benefit from using HiRA as part of the
evaluation process, as it gives a more equitable and appropriate reflection
of model performance at higher resolutions.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e144">When developing and improving forecast models, an important aspect is to
assess whether model changes have truly improved the forecast. Assessment
can be a mixture of subjective approaches, such as visualising forecasts and
assessing whether the broad structure of a field is appropriate, or
objective methods, comparing the difference between the forecast and an
observed or analysed value of “truth” for the model domain.</p>
      <p id="d1e147">Different types of intercomparison can be applied to identify the following different
underlying behaviours:
<list list-type="bullet"><list-item>
      <p id="d1e152">between different forecasting systems over an overlapping region to check
for model consistency between the two;</p></list-item><list-item>
      <p id="d1e156">between two versions of the same model to test the value of model upgrades
prior to operational implementation;</p></list-item><list-item>
      <p id="d1e160">parent–son intercomparison, evaluating the impact of downscaling or nesting
of models;</p></list-item><list-item>
      <?pagebreak page832?><p id="d1e164">a forecast comparison against reanalysis of the same model, inferring the
effect of resolution and forcing, especially in coastal areas.</p></list-item></list>
There are a number of works which have used these types of assessment to
delve into the characteristics of forecast models (e.g. Aznar et al.,
2016;
Mason et al., 2019; Juza et al., 2015) and produce coordinated validation
approaches (Hernandez et al., 2015).</p>
      <p id="d1e168">To aid the production of quality model assessment, services exist which
regularly produce multi-model assessments to deliver to the ocean community
(e.g. Lorente et al., 2019b).</p>
      <p id="d1e171">One of the issues faced when assessing high-resolution models against
lower-resolution models over the same domain is that often the coarser model
appears to perform at least equivalently or better when using typical
verification metrics such as root mean squared error (RMSE) or mean error,
which is a measure of the bias. Whereas a higher-resolution model has the
ability and requirement to forecast greater variation, detail and extremes,
a coarser model cannot resolve the detail and will, by its nature, produce
smoother features with less variation resulting in smaller errors. This can
lead to the situation that despite the higher-resolution model looking more
realistic it may verify worse (e.g. Mass et al., 2002; Tonani et
al., 2019).</p>
      <p id="d1e175">This is particularly the case when assessing forecast models categorically.
If the location of a feature in the model is incorrect, then two penalties
will be accrued: one for not forecasting the feature where it should have
been and one for forecasting the same feature where it did not occur (the
double-penalty effect, e.g. Rossa et al., 2008). This effect is more
prevalent in higher-resolution models due to their ability to, at least,
partially resolve smaller-scale features of interest. If the
lower-resolution model could not resolve the feature and therefore did not
forecast it, that model would only be penalised once. Therefore, despite
giving potentially better guidance, the higher-resolution model will verify
worse.</p>
      <p id="d1e178">Yet, the underlying need to quantitatively show the value of high-resolution
led to the development of so-called “spatial” verification methods which
aimed to account for the fact the forecast produced realistic features that
were not necessarily at the right place or at quite the right time (e.g. Ebert, 2008, or Gilleland, 2009). These methods have been in routine use
within the atmospheric model community for a number of years with some
long-term assessments and model comparisons (e.g. Mittermaier et al., 2013, for
precipitation).</p>
      <p id="d1e181">Spatial methods allow forecast models to be assessed with respect to several
different types of focus. Initially, these methods were classified into four
groups. Some methods look at the ability to forecast specific features (e.g. Davis et al., 2006); some look at how well the model performs at different
scales (scale separation, e.g. Casati et al., 2004). Others look at field
deformation (how much a field would have to be transformed to match a
“truth” field (e.g. Keil and Craig, 2007). Finally, there is neighbourhood
verification, many of which are equivalent to low-band-pass filters. In
these methods, forecasts are assessed at multiple spatial or temporal scales
to see how model skill changes as the scale is varied.</p>
      <p id="d1e184">Dorninger et al. (2018) provides an updated classification of spatial
methods, suggesting a fifth class of methods, known as distance metrics,
which sit between field deformation and feature-based methods. These methods
evaluate the distances between features, but instead of just calculating the
difference in object centroids (which is typical), the distances between all
grid point pairs are calculated, which makes distance metrics more similar
to field deformation approaches. Furthermore, there is no prior
identification of features. This makes distance metrics a distinct group
that warrants being treated as such in terms of classification. Not all
methods are easy to classify. An example of this is the integrated ice edge
error (IIEE) developed for assessing the sea ice extent (Goessling et al.,
2016).</p>
      <p id="d1e187">This paper exploits the use of one such spatial technique for the
verification of sea surface temperature (SST) in order to determine the
levels of forecast accuracy and skill across a range of model resolutions.
The High-Resolution Assessment framework (Mittermaier, 2014; Mittermaier and
Csima, 2017) is applied to the Met Office Atlantic Margin Model running at
7 <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> (O'Dea et al., 2012, 2017; King et al., 2018) (AMM7) and
1.5 <inline-formula><mml:math id="M2" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> (Graham et al., 2018; Tonani et al., 2019) (AMM15) resolutions for
the  North West European Shelf (NWS). The aim is to deliver an improved
understanding beyond the use of basic biases and RMSEs for assessing
higher-resolution ocean models, which would then better inform users on the
quality of regional forecast products. Atmospheric science has been using
high-resolution convective-scale models for over a decade and thus have
experience in assessing forecast skill on these scales, so it is appropriate
to trial these methods on eddy-resolving ocean model data. As part of the
demonstration, the paper also looks at how the method should be applied to
different ocean areas, where variation at different scales occurs due to
underlying driving processes.</p>
      <p id="d1e206">The paper was influenced by discussions on how to quantify the added value
from investments in higher-resolution modelling given the issues around the
double-penalty effect discussed above, which is currently an active area of
research within the ocean community (Lorente et al., 2019a; Hernandez et
al., 2018; Mourre et al., 2018).</p>
      <p id="d1e210">Section 2 describes the model and observations used in this study along with
the method applied. Section 3 presents the results, and Sect. 4 discusses
the lessons learnt while using HiRA on ocean forecasts and sets the path for
future work by detailing the potential and limitations of the method.</p>
</sec>
<?pagebreak page833?><sec id="Ch1.S2">
  <label>2</label><title>Data and methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Forecasts</title>
      <p id="d1e228">The forecast data used in this study are from the two products available in
the Copernicus Marine Environment Monitoring Service (CMEMS; see, e.g. Le
Traon et al., 2019, for a summary of the service) for the North West
European Shelf area:
<list list-type="bullet"><list-item>
      <p id="d1e233">NORTHWESTSHELF_ANALYSIS_FORECAST_ PHYS_004_001_b (AMM7) and</p></list-item><list-item>
      <p id="d1e237">NORTHWESTSHELF_ANALYSIS_FORECAST_ PHY_004_013 (AMM15).</p></list-item></list>
The major difference between these two products is the horizontal
resolution: <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M4" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> for AMM7 and 1.5 <inline-formula><mml:math id="M5" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> for AMM15. Both systems
are based on a forecasting ocean assimilation model with tides. The ocean
model is NEMO (Nucleus for European Modelling of the Ocean; Madec and the NEMO team, 2016),
using the 3D-Var NEMOVAR system to assimilate observations (Mogensen et al.,
2012). These are surface temperature in situ and satellite measurements,
vertical profiles of temperature and salinity, and along-track satellite sea
level anomaly data. The models are forced by lateral boundary conditions
from the UK Met Office North Atlantic Ocean forecast model and by the CMEMS
Baltic forecast product (BALTICSEA_ANALYSIS_FORECAST_PHY_003_006). The
atmospheric forcing is given by the operational European Centre for
Medium-Range Weather Forecasts (ECMWF) numerical weather prediction model
for AMM15 and by the operational UK Met Office Global Atmospheric model for
AMM7.</p>
      <p id="d1e267">The AMM15 and AMM7 systems run once a day and provide forecasts for
temperature, salinity, horizontal currents, sea level, mixed layer depth
and bottom temperature. Hourly instantaneous values and daily 25 <inline-formula><mml:math id="M6" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">h</mml:mi></mml:mrow></mml:math></inline-formula>
de-tided averages are provided for the full water column.</p>
      <p id="d1e278">AMM7 has a regular latitude–longitude grid, whilst AMM15 is computed on a
rotated grid and regridded to have both models delivered to the (CMEMS)
data catalogue (<uri>http://marine.copernicus.eu/services-portfolio/access-to-products/</uri>,
last access: October 2019) on a
regular grid. Table 1 provides a summary of the model configuration, a fuller description can be found in Tonani et al. (2019).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e287">AMM7 and AMM15 co-located areas. Note the difference in the
land–sea boundaries due to the different resolutions, notably around the
Scandinavian coast. Contours show the model bathymetry at 200, 2000 and 4000 <inline-formula><mml:math id="M7" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f01.png"/>

        </fig>

      <p id="d1e304">For the purpose of this assessment, the 5 <inline-formula><mml:math id="M8" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:math></inline-formula> daily mean potential SST
forecasts (with lead times of 12, 36, 60, 84 and
108 <inline-formula><mml:math id="M9" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">h</mml:mi></mml:mrow></mml:math></inline-formula>) were utilised for the period from January to September 2019.
Forecasts were compared for the co-located areas of AMM7 and AMM15. Figure 1
shows the AMM7 and AMM15 co-located domain along with the land–sea mask for
each of the models. AMM15 has a more detailed coastline and SST field than
AMM7 due to its higher resolution. When comparing two models with different
resolutions, it is important to know whether increased detail actually
translates into better forecast skill. Additionally, the differences in
coastline representation can have an impact on any HiRA results obtained, as
will be discussed in a later section.</p>
      <p id="d1e323">It should be noted that this study is an assessment of the application of
spatial methods to ocean forecast data and, as such, is not meant as a full
and formal assessment and evaluation of the forecast skill of the AMM7 and
AMM15 ocean configurations. To this end, a number of considerations have had
to be taken into account in order to reduce the complexity of this initial
study. Specifically, it was decided at an early stage to use daily mean SST
temperatures, as opposed to hourly instantaneous SST, as this avoided any
influence of the diurnal cycle and tides on any conclusions made. AMM15 and
AMM7 daily means are calculated as means over 25 <inline-formula><mml:math id="M10" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">h</mml:mi></mml:mrow></mml:math></inline-formula> to remove both the
diurnal cycle and the tides. The tidal signal is removed because the period
of the major tidal constituent, the semidiurnal lunar component M2, is 12 <inline-formula><mml:math id="M11" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">h</mml:mi></mml:mrow></mml:math></inline-formula>
and 25 <inline-formula><mml:math id="M12" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">min</mml:mi></mml:mrow></mml:math></inline-formula> (Howarth and Pugh, 1983). Daily means are also one of the
variables that are available from the majority of the products within the
CMEMS catalogue, including reanalysis, so the application of the spatial
methods could be relevant in other use cases beyond those considered here.
In addition, there are differences in both the source and frequency of the
air–sea interface forcing used in both the AMM7 and AMM15 configurations
which could influence the results. Most notably, AMM7 uses hourly
surface pressure and 10 <inline-formula><mml:math id="M13" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> winds from the Met Office Unified Model (UM),
whereas AMM15 uses 3-hourly data from ECMWF.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e360">Observation locations within the domain for 12:00 UTC on 6 June 2019.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f02.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Observations</title>
      <p id="d1e377">SST observations used in the verification were downloaded from the CMEMS
catalogue from the product
<list list-type="bullet"><list-item>
      <p id="d1e382">INSITU_NWS_NRT_OBSERVATIONS_013_036.</p></list-item></list>
This dataset consists of in situ observations only, including daily
drifters, mooring, ferry-box and conductivity–temperature–depth (CTD)
observations. This results in a varying number of observations being
available throughout the verification period, with uneven spatial coverage
over the verification domain. Figure 2 shows a
snapshot of the typical observational coverage, in this case for 12:00 UTC on
6 June 2019. This coverage is important when assessing the results,
notably when thinking about the size and type of area over which an
observation is meant to be representative of, and how close to the coastline
each observation is.</p>
      <p id="d1e386">This study was set up to detect issues that should be considered by users
when applying HiRA within a routine ocean verification setup, using a broad
assessment containing as much data as were available in order to understand
the impact of using HiRA for ocean forecasts. Several assumptions were made
in this study.</p>
      <p id="d1e389">For example, there is a temporal mismatch between the forecasts and
observations used. The forecasts (which were available at the time of this
study) are daily means of the SSTs from 00:00 to 00:00 UTC, whilst the
observations are<?pagebreak page834?> instantaneous and usually available hourly. For the
purposes of this assessment, we have focused on SSTs closest to the
midpoint of the forecast period for each day (nominally 12:00 UTC).
Observation times had to be within 90 <inline-formula><mml:math id="M14" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">min</mml:mi></mml:mrow></mml:math></inline-formula> of this time, with any other
times from the same observation site being rejected. A particular reason for
picking a single observation time rather than daily averages was so that
moving observations, such as drifting buoys, could be incorporated into the
assessment. Creating daily mean observations from moving observations would
involve averaging reports from different forecast grid boxes and hence
contaminate the signal that HiRA is trying to evaluate.</p>
      <p id="d1e400">Future applications would probably contain a stricter setup, e.g. only
using fixed daily mean observations, or verifying instantaneous (hourly)
forecasts so as to provide a subdaily assessment of the variable in
question.</p><?xmltex \hack{\newpage}?>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>HiRA</title>
      <p id="d1e413">The HiRA framework (Mittermaier, 2014) was designed to overcome the
difficulties encountered in assessing the skill of high-resolution models
when evaluating against point observations. Traditional verification metrics
such as RMSE and mean error rely on precise matching in space and time, by
(typically) extracting the nearest model grid point to an observing
location. The method is an example of a
single-observation neighbourhood-forecast (SO-NF) approach, with no
smoothing. All the forecast grid points within a neighbourhood centred on an
observing location are treated as a pseudo ensemble, which is evaluated
using well-known ensemble and probabilistic forecast metrics. Scores are
computed for a range of (increasing) neighbourhood sizes to understand the
scale–error relationship. This approach assumes that the observation is
not only representative of  its precise location but also has characteristics
of the surrounding area. WMO manual No. 8 (World Meteorological Organisation, 2017) suggests that, in
the atmosphere, observations can be considered representative of an
area within a 100 <inline-formula><mml:math id="M15" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> radius of a land station, but this is often very
optimistic. The manual states further: “For small-scale or local
applications the considered area may have dimensions of 10 km or less.” A
similar principle applies to the ocean; i.e. observations can represent an
area around the nominal observation location, though the representative
scales are likely to be very different from in the atmosphere. The
representative scale for an observation will also depend on local
characteristics of the area, for example, whether the observation is on the
shelf or in the open ocean or likely to be impacted by river discharge.</p>
      <p id="d1e424">There will be a limit to the useful forecast neighbourhood size which can be
used when comparing to a point observation. This maximum neighbourhood size
will depend on the representative scale of the variable under consideration.
Put differently, once the neighbourhoods become too big, there will be
forecast values in the pseudo ensemble which will<?pagebreak page835?> not be representative of
the observation (and the local climatology) and any skill calculated will be
essentially random. Combining results for multiple observations with very
different representative scales (for example, a mixture of deep ocean and
coastal observations) could contaminate results due to the forecast
neighbourhood only being representative of a subset of the observations. The
effect of this is explored later in this paper.</p>
      <p id="d1e427">HiRA can be based on a range of statistics, data thresholds and
neighbourhood sizes in order to assess a forecast model. When comparing
deterministic models of different resolutions, the approach is to equalise
on the physical area of the neighbourhoods (i.e. having the same
“footprint”). By choosing sequences of neighbourhoods that provide (at
least) approximate equivalent neighbourhoods (in terms of area), two or more
models can be fairly compared.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e433">Example of forecast grid point selections for different HiRA
neighbourhoods for a single observation point. A <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> domain returns nine points
that represent the nearest forecast grid points in a square around the
observation. A <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> domain encompasses more points.</p></caption>
        <?xmltex \igopts{width=142.26378pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f03.png"/>

      </fig>

      <p id="d1e466">HiRA works as follows. For each observation, several neighbourhood sizes are
constructed, representing the length in forecast grid points of a square
domain around the observation points, centred on the grid point closest to
the observation (Fig. 3). There is no interpolation applied to the forecast
data to bring them to the observation point; all the data values are used
unaltered.</p>
      <p id="d1e469">Once neighbourhoods have been constructed, the data can be assessed using a
range of well-known ensemble or probabilistic scores. The choice of
statistic usually depends on the characteristics of the parameter being
assessed. Parameters with significant thresholds can be assessed using the
Brier score (Brier, 1950) or the ranked probability score (RPS) (Epstein,
1969), i.e. assessing the ability of the forecast to correctly locate a
forecast in the correct threshold band. For continuous variables such as
SST, the data have been assessed using the continuous ranked probability
score (CRPS) (Brown, 1974; Hersbach, 2000).</p>
      <p id="d1e472"><?xmltex \hack{\newpage}?>The CRPS is a continuous extension of the RPS. Whereas the RPS is
effectively an average of a user-defined set of Brier scores over a finite
number of thresholds, the CRPS extends this by considering an integral over
all possible thresholds. It lends itself well to ensemble forecasts of
continuous variables such as temperature and has the useful property that
the score reduces to the mean absolute error (MAE) for a single-grid-point
deterministic model comparison. This means that, if required, both
deterministic and probabilistic forecasts can be compared using the same
score.
          <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M18" display="block"><mml:mrow><mml:mtext>CRPS</mml:mtext><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi mathvariant="normal">∞</mml:mi></mml:munderover><mml:msup><mml:mfenced open="[" close="]"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">fcst</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mi>x</mml:mi></mml:mfenced><mml:mo>-</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">obs</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mi>x</mml:mi></mml:mfenced></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mi mathvariant="normal">d</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></disp-formula>
        Equation (1) defines the CRPS, where for a parameter <inline-formula><mml:math id="M19" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">fcst</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the
cumulative distribution of the neighbourhood forecast and <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">obs</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is
the cumulative distribution of the observed value, represented by a
Heaviside function (see Hersbach, 2000). The CRPS is an error-based score
where a perfect forecast has a value of zero. It measures the difference
between two cumulative distributions: a forecast distribution formed by
ranking the (in this case, quasi)-ensemble members represented by the
forecast values in the neighbourhood and a step function describing the
observed state. To use an ensemble, HiRA makes the assumption that all grid
points within a neighbourhood are equi-probable outcomes at the observing
location. Therefore, aside from the observation representativeness limit, as
the neighbourhood sizes increase, this assumption of equi-probability will
break down as well, and scores become random. Care must therefore be taken
to decide whether a particular neighbourhood size is appropriately
representative. This decision will be based on the length scales appropriate
for a variable as well as the resolution of the forecast model being
assessed. Figure 4 shows a schematic of how
different neighbourhood sizes contribute towards constructing forecast
probability density functions around a single observation.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e566">Example of how different forecast neighbourhood sizes would
contribute to the generation of a probability density function around an
observation (denoted by x). The larger the neighbourhood, the better
described the pdf, though potentially at the expense of larger spread. If
a forecast point is invalid within the forecast neighbourhood, that site
is rejected from the calculations for that neighbourhood size.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f04.png"/>

      </fig>

      <?pagebreak page836?><p id="d1e576"><?xmltex \hack{\newpage}?>AMM7 and AMM15 resolve different length scales of motion due to their
horizontal resolution. This should be taken into account when assessing the
results of different neighbourhood sizes. Both models can resolve the large
barotropic scale (<inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M23" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>) and the shorter baroclinic scale
off the shelf in deep water. On the continental shelf, only the resolution
of <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M25" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> of AMM15 permits motions at the smallest
baroclinic scale since the first baroclinic Rossby radius is on the order of 4 <inline-formula><mml:math id="M26" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>
(O'Dea et al., 2012). AMM15 represents a step change in representing the
eddy dynamics variability on the continental shelf. This difference has an
impact also on the data assimilation scheme, where two horizontal
correlation length scales (Mirouze et al., 2016) are used to represent large
and small scales of ocean variability. The long length scale is 100 <inline-formula><mml:math id="M27" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>, while
the short correlation length scale aims to account for internal ocean
processes variability, characterised by the Rossby radius of deformation.
Computational requirements restrict the short length scale to be at least three
model grid points, 4.5 and 21 <inline-formula><mml:math id="M28" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>, respectively, for AMM15 and AMM7 (Tonani
et al., 2019). Although AMM15 resolves smaller-scale processes, comparing
AMM7 and AMM15 in neighbourhood sizes between the AMM7 resolution and
multiples of this resolution will address processes that should be accounted
for in both models.</p>
      <p id="d1e641">As the methodology is based on ensemble and probabilistic metrics, it is
naturally extensible to ensemble forecasts (see Mittermaier and Csima,
2017), which are currently being developed in research mode by the ocean
community, allowing for intercomparison between deterministic and
probabilistic forecast models in an equitable and consistent way.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Model Evaluation Tools (MET)</title>
      <p id="d1e652">Verification was performed using the Point-Stat tool, which is part of the
Model Evaluation Tools (MET) verification package that was developed by the
National Center for Atmospheric Research (NCAR) and which can be configured
to generate CRPS results using the HiRA framework. MET is free to download
from GitHub at <uri>https://github.com/NCAR/MET</uri> (last access: May 2019).</p>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Equivalent neighbourhoods and equalisation</title>
      <p id="d1e667">When comparing neighbourhoods between models, the preference is to look for
similar-sized areas around an observation and then transform this to the
closest odd-numbered square neighbourhood, which will be called the
“equivalent neighbourhood”. In the case of the two models used, the most
appropriate neighbourhood size can change depending on the structure of the
grid, so the user needs to take into consideration what is an accurate match
between the models being compared.</p>
      <p id="d1e670">The two model configurations used in this assessment are provided on
standard latitude–longitude grids via the CMEMS catalogue. The AMM7 and
AMM15 configurations are stated to have resolutions approximating 7 and
1.5 <inline-formula><mml:math id="M29" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>, respectively. Thus, equivalent neighbourhoods should simply be a case
of matching neighbourhoods with similar spatial distances. In fact,
AMM15 is originally run on a rotated latitude–longitude grid where the
resolution is closely approximated by 1.5 <inline-formula><mml:math id="M30" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> and subsequently provided to
the CMEMS catalogue on the standard latitude–longitude grid. Once the grid
has been transformed to a regular latitude–longitude grid, the 1.5 <inline-formula><mml:math id="M31" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> nominal
spatial resolution is not as accurate. This is particularly important when
neighbourhood sizes become larger, since any error in the approximation of
the resolution will become multiplied as the number of points being used
increases.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e700">Summary of the main differences between
NORTHWESTSHELF_ANALYSIS_FORECAST_PHYS_004_001_b (AMM7) and
NORTHWESTSHELF_ANALYSIS_FORECAST_PHYS_004_013 (AMM15).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Resolution</oasis:entry>
         <oasis:entry colname="col3">Atmospheric forcing</oasis:entry>
         <oasis:entry colname="col4">Geographical model domain</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">AMM7</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M33" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">MetUM 10 <inline-formula><mml:math id="M34" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">40–65<inline-formula><mml:math id="M35" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 20<inline-formula><mml:math id="M36" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W–13<inline-formula><mml:math id="M37" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">AMM15</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M39" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">ECMWF IFS <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M41" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">45</mml:mn></mml:mrow></mml:math></inline-formula>–63<inline-formula><mml:math id="M43" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M45" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W–13<inline-formula><mml:math id="M46" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e899">Details of equivalent neighbourhoods used when comparing AMM7 and
AMM15.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry rowsep="1" namest="col2" nameend="col5" align="center" colsep="1">AMM7 </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col9" align="center">AMM15 </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Name</oasis:entry>
         <oasis:entry colname="col2">Total points</oasis:entry>
         <oasis:entry colname="col3">Shape</oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">Size (E–W) </oasis:entry>
         <oasis:entry colname="col6">Total points</oasis:entry>
         <oasis:entry colname="col7">Shape</oasis:entry>
         <oasis:entry rowsep="1" namest="col8" nameend="col9" align="center">Size (E–W) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">Actual (<inline-formula><mml:math id="M47" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col5">Nominal (km)</oasis:entry>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8">Actual (<inline-formula><mml:math id="M48" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col9">Nominal (km)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">NB1</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.11</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">25</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.15</oasis:entry>
         <oasis:entry colname="col9">7.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NB2</oasis:entry>
         <oasis:entry colname="col2">9</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.33</oasis:entry>
         <oasis:entry colname="col5">21</oasis:entry>
         <oasis:entry colname="col6">121</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:mn mathvariant="normal">11</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.33</oasis:entry>
         <oasis:entry colname="col9">16.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NB3</oasis:entry>
         <oasis:entry colname="col2">25</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.55</oasis:entry>
         <oasis:entry colname="col5">35</oasis:entry>
         <oasis:entry colname="col6">361</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mn mathvariant="normal">19</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">19</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.57</oasis:entry>
         <oasis:entry colname="col9">28.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NB4</oasis:entry>
         <oasis:entry colname="col2">49</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mn mathvariant="normal">7</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.77</oasis:entry>
         <oasis:entry colname="col5">49</oasis:entry>
         <oasis:entry colname="col6">625</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mn mathvariant="normal">25</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">25</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.76</oasis:entry>
         <oasis:entry colname="col9">37.5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NB5</oasis:entry>
         <oasis:entry colname="col2">81</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mn mathvariant="normal">9</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">9</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">0.99</oasis:entry>
         <oasis:entry colname="col5">63</oasis:entry>
         <oasis:entry colname="col6">1089</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mn mathvariant="normal">33</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">33</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.99</oasis:entry>
         <oasis:entry colname="col9">49.5</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e1268">Additionally, the two model configurations do not have the same aspect ratio
of grid points. AMM7 has a longitudinal resolution of <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.11</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M60" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> and a latitudinal resolution of <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.066</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M62" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> (a ratio of <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>), whilst the AMM15 grid has a resolution of
<inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mo>∼</mml:mo><mml:mn mathvariant="normal">0.0135</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M66" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>,
respectively (a ratio of <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:math></inline-formula>). HiRA neighbourhoods typically contain the
same number of grid points in the zonal and meridional directions which will
lead to discrepancies in the area selected when comparing models with
different grid aspect ratios, depending on whether the comparison is based
on neighbourhoods with a similar longitudinal or similar latitudinal size.
This difference will scale as the neighbourhood size increases, as shown in
Fig. 4 and Table 2. The onus is therefore on the user to understand any
difference in grid structure, and therefore within the HiRA neighbourhoods,
between models being compared and to allow for this when comparing
equivalent neighbourhoods.</p>
      <p id="d1e1360">For this study, we have matched neighbourhoods between model configurations
based on their longitudinal size. The equivalent neighbourhoods used to show
similar areas within the two configurations are indicated in Table 2 along
with the bar style and naming convention used throughout.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e1365">Similar neighbourhood sizes for a 49 <inline-formula><mml:math id="M68" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> neighbourhood using the
approximate resolutions (7 and 1.5 <inline-formula><mml:math id="M69" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>) with <bold>(a)</bold> AMM7 with a <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mn mathvariant="normal">7</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:math></inline-formula>
neighbourhood and <bold>(b)</bold> AMM15 with a <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mn mathvariant="normal">33</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">33</mml:mn></mml:mrow></mml:math></inline-formula> neighbourhood. Whilst the
neighbourhoods are similar sizes in the latitudinal direction, the AMM15
neighbourhood is sampling a much larger area due to different scales in the
longitudinal direction. This means that a comparison with a <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mn mathvariant="normal">25</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">25</mml:mn></mml:mrow></mml:math></inline-formula> AMM15
neighbourhood is more appropriate.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f05.png"/>

      </fig>

      <p id="d1e1433">For ocean applications, there are other aspects of the processing to be aware
of when using neighbourhood methods. This is mainly related to the presence
of coastlines and<?pagebreak page837?> how their representation changes resolution (as defined by
the land–sea mask) and the treatment of observations within HiRA
neighbourhoods. Figure 5 illustrates the contrasting land–sea boundaries due
to the different resolutions of the two configurations. When calculating
HiRA neighbourhood values, all forecast values in the specific neighbourhood
around an observation must be present for a score to be calculated. If any
forecast points within a neighbourhood contain missing data, then that
observation at that neighbourhood size is rejected. This is to ensure that
the resolution of the “ensemble”, which is defined or determined by the
number of members, remains the same. For typical atmospheric fields such as
screen temperature, this is not an issue, but with parameters that have
physical boundaries (coastlines), such as SST, there will be discontinuities
in the forecast field that depend on the location of the land–sea boundary.
For coastal observations, this means that as the neighbourhood size
increases, it is more likely that an observation will be rejected from the
comparison due to missing data. Even at the grid scale, the nearest model
grid point to an observation may not be a sea point. In addition, different
land–sea borders between models mean that potentially some observations will
be rejected from one model comparison but will be retained in the other
because of missing forecast points within their respective neighbourhoods.
Care should be taken when implementing HiRA to check the observations
available to each model configuration when assessing the results and make a
judgement as to whether the differences are important.</p>
      <p id="d1e1437">There are potential ways to ensure equalisation, for example, only using
observations that are available in both configurations for a location and
neighbourhoods, or only observations away from the coast. For the purposes
of this study, which aims to show the utility of the method, it was judged
important to use as many observations as possible, so as to capture any
potential pitfalls in the application of the framework, which would be
relevant to any future application of it.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e1442">Number of observation sites within NB1, NB3 and NB5 for AMM15 and
AMM7. Numbers are those used during September 2019 but represent typical
total observations during a month. Matching line styles represent equivalent
neighbourhoods.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f06.png"/>

      </fig>

      <p id="d1e1451">Figure 6 shows the number of observations available
to each neighbourhood for each day during September 2019. For each model
configuration, it shows how these observations vary within the HiRA
framework. There are several reasons for the differences shown in the plot.
There is the difference mentioned previously whereby a model neighbourhood
includes a land point and therefore is rejected from the calculations
because the number of quasi-ensemble members is no longer the same. This is
more likely for coastal observations and depends on the particularities of
the model land–sea mask near each observation. This rejection is more likely
for the high-resolution AMM15 when looking at equivalent areas, in part due
to the larger number of grid boxes being used; however, there are also
instances of observations being rejected from the coarser-resolution AMM7
and not the higher-resolution AMM15 due to nuances of the land–sea mask.</p>
      <p id="d1e1454">It is apparent that for equivalent neighbourhoods there are typically more
observations available for the coarser model configuration and that this
difference is largest for the smallest equivalent neighbourhood size but
becoming less obvious at larger neighbourhoods. It could therefore be worth
considering that the large benefit in AMM15 when looking at the first
equivalent neighbourhood is potentially influenced by the difference in
observations. As the neighbourhood sizes increase, the number of
observations reduces due to the higher likelihood of a land point being part
of a larger neighbourhood. It is also noted that there is a general daily
variability in the number of observations present, based on differences in
the observations reporting on any particular day within the co-located
domain.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e1459">Verification results using a typical statistics approach for
January–September 2019. Mean error <bold>(a)</bold>, root mean square error <bold>(b)</bold>
and mean absolute error <bold>(c)</bold> results are shown for the two model
configurations. Two methods of matching forecast to observations points have
been used: a nearest-neighbour approach (solid) representing the
single-grid-point results from HiRA and a bilinear interpolation approach (dashed) more
typically used in operational ocean verification.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f07.png"/>

      </fig>

</sec>
<?pagebreak page838?><sec id="Ch1.S6">
  <label>6</label><title>Results</title>
      <p id="d1e1485">Figure 7 shows the aggregated results from the study
period defined in Sect. 2 by applying typical verification statistics.
Results have been averaged across the entire period from January to
September and output relative to the forecast validity time. Two methods of
matching forecast grid points to observation locations have been used.
Bilinear interpolation is typically the approach used in traditional
verification of SST, as it is a smoothly varying field. A nearest-neighbour
approach has also been shown, as this is the method that would be used for
HiRA when applying it at the grid scale.</p>
      <p id="d1e1488">It is noted that the two methods of matching forecasts to observation
locations give quite different results. For the mean error, the impact of
moving from a single-grid-point approach to a bilinear interpolation method
appears to be minor for the AMM7 model but is more severe for AMM15,
resulting in a larger error across all lead times. For the RMSE, the picture
is more mixed, generally suggesting that the AMM7 forecasts are better when
using a bilinear interpolation method but giving no clear overall steer when
the nearest grid point is used. However, the impact of taking a bilinear
approach results in much higher gross errors across all lead times when
compared to the nearest grid point approach.</p>
      <p id="d1e1491">The MAE has been suggested as a more appropriate metric than the RMSE for
ocean fields using (as is the case here) near-real-time observation data
(Brassington, 2017). In Fig. 6, it can be seen that the nearest grid point
approach for both AMM7 and AMM15 gives almost exactly the same results,
except for the shortest of lead times. For the bilinear interpolation
method, AMM15 has a smaller error than AMM7 as lead time increases, behaviour
which is not apparent when RMSE is applied.</p>
      <p id="d1e1494">Based on the interpolated RMSE results in Fig. 6, it would be hard to
conclude that there was a significant benefit to using high-resolution ocean
models for forecasting SSTs. This is where the HiRA framework can be
applied. It can be used<?pagebreak page839?> to provide more information, which can better inform
any conclusions on model error.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e1500">Summary of CRPS (left axis, lines) and CRPS difference (right
axis, bars) for the January 2019 to September 2019  period  for AMM7 and AMM15
models at different neighbourhood sizes. Error bars represent 95 %
confidence intervals generated using a bootstrap with replacement method for
10 000 samples. An “S” above the bar denotes that 95 % error bars for the
two models do not overlap.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f08.png"/>

      </fig>

      <p id="d1e1509">Figure 8 shows the results for AMM7 and AMM15 for the January–September 2019
period using the HiRA framework with the CRPS. The lines on the plot
show the CRPS for the two model configurations for different neighbourhood
sizes, each plotted against lead time. Similar line styles are used to
represent equivalent neighbourhood sizes. Confidence intervals have been
generated by applying a bootstrap with replacement method, using 10 000
samples, to the domain-averaged CRPS (e.g. Efron and Tibshirani, 1986). The
error bars represent the 95 % confidence level. The results for the
single grid point show the MAE and are the same as would be obtained using a
traditional (precise) matching. In the case of CRPS, where a lower score is
better, we see that AMM15 is better than AMM7, though not significantly so,
except at shorter lead times where there is little difference.</p>
      <p id="d1e1512">The differences at equivalent neighbourhood sizes are displayed as a bar
plot on the same figure, with scores referenced with respect to the
right-hand axis. Line markers and error bars have been offset to aid
visualisation, such that results for equivalent neighbourhoods are displayed
in the same vertical column as the difference indicated by the bar plot. The
details of the equivalent neighbourhood sizes are presented in Table 2.
Since a lower CRPS score is better, a positively orientated (upwards) bar
implies AMM7 is better, whilst a negatively orientated (downwards) bar means
AMM15 is better.</p>
      <p id="d1e1515">As indicated in Table 2, NB1 compares the single-grid-point results of AMM7
with a 25-member pseudo-ensemble constructed from a <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> AMM15 neighbourhood.
Given the different resolutions of the two configurations, these two
neighbourhoods represent similar physical areas from each model domain, with
AMM7 only represented by a single forecast value for each observation but
AMM15 represented by 25 values covering the same area, and as such potentially
better able to represent small-scale variability within that area.</p>
      <p id="d1e1530">At this equivalent scale, the AMM15 results are markedly better than AMM7,
with lower errors, suggesting that overall the AMM15 neighbourhood better
represents the variation around the observation than the coarser single grid
point of AMM7. In the next set of equivalent neighbourhoods (NB2), the gap
between the two configurations has closed, but AMM15 is still consistently
better than AMM7 as lead time increases. Above this scale, the neighbourhood
values tend towards similarity and then start to diverge again suggesting
that the representative scale of the neighbourhoods has been reached and
that errors are essentially random.</p>
      <p id="d1e1533">Whilst the overall HiRA neighbourhood results for the co-located domains
appear to show a benefit to using a higher-resolution model forecast, it
could be that these results are influenced by the spatial distribution of
observations within the domain and the characteristics of the forecasts at
those locations. In order to investigate whether this was important
behaviour, the results were separated into two domains: one representing the
continental shelf part of the domain (where the bathymetry <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M75" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>)
and the other representing the deeper, off-shelf, ocean component (Fig. 9).
HiRA results were compared for observations only within each masked domain.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9"><?xmltex \currentcnt{9}?><label>Figure 9</label><caption><p id="d1e1557">On-shelf and off-shelf masking regions within the co-located AMM7
and AMM15 domain (data within the grey areas are masked).</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f09.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><label>Figure 10</label><caption><p id="d1e1568">Summary of on-shelf CRPS (left axis, lines) and CRPS difference
(right axis, bars) for the January 2019 to September 2019 period  for AMM7
and AMM15 models at different neighbourhood sizes. Error bars represent 95 % confidence values obtained from 10 000 samples using bootstrap with
replacement. An “S” above the bar denotes that 95 % error bars for the
two models do not overlap.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f10.png"/>

      </fig>

      <p id="d1e1577">On-shelf results (Fig. 10) show that at the grid scale the results for both
AMM7 and AMM15 are worse for this subdomain. This could be explained by
both the complexity of processes (tides, friction, river mixing,
topographical effects, etc.) and the small dynamical scales associated with
shallow waters on the shelf (Holt et al., 2017).</p>
      <p id="d1e1580">The on-shelf spatial variability in SST across a neighbourhood is likely to
be higher than for an equivalent deep ocean neighbourhood due to small-scale
changes in bathymetry and, for some observations, the impact of coastal
effects. Both AMM7 and AMM15 show improvement in CRPS with increased
neighbourhood size until the CRPS plateaus in the range 0.225 to 0.25, with
AMM15 generally better than AMM7 for equivalent neighbourhood sizes. Scores
get worse (errors increase) for both model configurations as the forecast
lead time increases.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><label>Figure 11</label><caption><p id="d1e1585">Summary of off-shelf CRPS (left axis, lines) and CRPS
difference (right axis, bars) for the January 2019 to September 2019  period
for AMM7 and AMM15 models at different neighbourhood sizes. Error bars
represent 95 % confidence values obtained from 10 000 samples using
bootstrap with replacement. An “S” above the bar denotes that 95 % error
bars for the two models do not overlap.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f11.png"/>

      </fig>

      <p id="d1e1595">For off-shelf results (Fig. 11), CRPS is much better (smaller error) at
both the grid scale and for HiRA neighbourhoods, suggesting that both
configurations are better at forecasting these deep ocean SSTs (or that it
is easier to do so). There is still an improvement in CRPS when going from
the grid scale (single grid box) to neighbourhoods, but the value of that
change is much smaller than that for the on-shelf subdomain. When comparing
equivalent neighbourhoods, AMM15 still gives consistently better results
(smaller errors) and appears to improve over AMM7 as lead time increases in
contrast to the on-shelf results.</p>
      <p id="d1e1598">It is likely that the neighbourhood at which we lose representativity will
be larger for the deeper ocean than the shelf area because of the larger
scale of dynamical processes in deep water. When choosing an optimum
neighbourhood to use for assessment, care should be taken to check whether
there are different representativity levels in the data (such as here for
on-shelf and off-shelf) and pragmatically choose the smaller of those
equivalent neighbourhoods when looking at data combining the different
representativity levels.</p>
      <p id="d1e1601">Overall, for the  January–September 2019 period, AMM15 demonstrates a
lower (better) CRPS than AMM7 when looking at the HiRA neighbourhoods.
However, this also appears to be true at the grid scale over the assessment
period. One of the aspects that HiRA is trying to provide additional
information about is whether higher-resolution models can demonstrate
improvement over coarser models against a perception that the coarser models
score better in standard verification forecast assessments. Assessed over
the whole period, this initial premise does not appear to hold true;
therefore, a deeper look at the data is required to assess whether this
signal is consistent within shorter time periods or if there<?pagebreak page840?> are
underlying periods contributing significant and contrasting results to the
whole-period aggregate.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12"><?xmltex \currentcnt{12}?><label>Figure 12</label><caption><p id="d1e1606">Monthly time series of whole-domain CRPS scores for the grid scale
(solid line) and NB2 neighbourhood (dashes) for <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">60</mml:mn></mml:mrow></mml:math></inline-formula> forecasts. Error bars
represent 95 % confidence values obtained from 10 000 samples using
bootstrap with replacement. Error bars have been staggered in the
<inline-formula><mml:math id="M77" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> direction to aid clarity.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f12.png"/>

      </fig>

      <p id="d1e1634">Figure 12 shows a monthly breakdown of the grid
scale and the NB2 HiRA neighbourhood scores at <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">60</mml:mn></mml:mrow></mml:math></inline-formula>. This shows the
underlying monthly variability not immediately apparent in the whole-period
plots. Notably for the January to March period, AMM7 outperforms AMM15 at
the grid scale. With the introduction of HiRA neighbourhoods, AMM7 still
performs better for February and March but the difference between the models
is significantly reduced. For these monthly time series, the error bars
increase in size relative to the summary plots (e.g. Fig. 8) due to the
reduction in data available. The sample size will have an impact on the
error bars as the smaller the sample, the less representative of the true
population the data are likely to be. April in particular contained several
days of missing forecast data, leading to a reduction in sample size and
corresponding increase in error bar size, whilst during May there was a
period with reduced numbers of observations.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13"><?xmltex \currentcnt{13}?><label>Figure 13</label><caption><p id="d1e1652">On-shelf monthly time series of CRPS. Error bars represent
95 % confidence values obtained from 10 000 samples using bootstrap with
replacement.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f13.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14"><?xmltex \currentcnt{14}?><label>Figure 14</label><caption><p id="d1e1663">Off-shelf monthly time series of CRPS. Error bars represent
95 % confidence values obtained from 10 000 samples using bootstrap with
replacement.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f14.png"/>

      </fig>

      <p id="d1e1672">The same pattern is present for the on-shelf subdomain (Fig. 13), where
what appears to be a significant benefit for AMM7 during February and
March is less clear-cut at the NB2 neighbourhood. For the off-shelf
subdomain (Fig. 14), differences between the two configurations at the grid
scale are mainly apparent during the summer months. At the NB2 scale,
AMM15 potentially demonstrates more benefit than AMM7 except for April and
May, where the two show similar results. There is a balance to be struck in
this conclusion as the differences between the two models are rarely greater
than the 95 % error bars. This in itself does not mean that the results
are not significant. However, care should be taken when interpreting such a
result as a statistical conclusion rather than broad guidance as to model
performance. Attempts to reduce the error bar size, such as increasing the
number of observations or number of times within the period, would aid this
interpretation.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F15"><?xmltex \currentcnt{15}?><label>Figure 15</label><caption><p id="d1e1677">Number of grid-scale observations for the on- and off-shelf
domains.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://os.copernicus.org/articles/16/831/2020/os-16-831-2020-f15.png"/>

      </fig>

      <p id="d1e1686">One noticeable aspect of the time series plots is that the whole-domain plot
is heavily influenced by the on-shelf results. This is due to the difference
in observation numbers as shown in Fig. 15, with the on-shelf domain having
more observations overall, sometimes significantly more, for example, during
January or mid-to-late August. For the overall domain, the on-shelf
observations will contribute more to the overall score, and hence the
underlying off-shelf signal will tend to be masked. This is an indication of
why verification is more useful when done over smaller, more homogeneous
subregions, rather than verifying everything together, with the caveat that
sample sizes are large enough, since underlying signals can be swamped by
dominant error types.</p>
</sec>
<sec id="Ch1.S7" sec-type="conclusions">
  <label>7</label><title>Discussion and conclusions</title>
      <p id="d1e1697">In this study, the HiRA framework has been applied to SST forecasts from two
ocean models with different resolutions. This enables a different view of
the forecast errors than<?pagebreak page841?> obtained using traditional (precise) grid-scale
matching against ocean observations. Particularly, it enables us to
demonstrate the additional value of high-resolution model. When considered
more appropriately, high-resolution models (with the ability to forecast
small-scale detail) have lower errors when compared to the smoother
forecasts provided by a coarser-resolution model.</p>
      <p id="d1e1700">The HiRA framework was intended to address the question of whether moving to
higher resolution adds value. This study has identified and highlighted
aspects that need to be considered when setting up such an assessment. Prior
to this study, routine verification statistics typically showed that
coarser-resolution models had equivalent skill to or more skill than higher-resolution models
(e.g. Mass et al., 2002; Tonani et al., 2019). During the
January to September 2019 period, grid-scale verification within this assessment
showed that the coarser-resolution AMM7 often demonstrated lower errors than
AMM15.</p>
      <p id="d1e1703">HiRA neighbourhoods were applied and the data then assessed using the CRPS,
showing a large reduction (improvement) in errors for AMM15 when going from
a grid-scale, point-based verification assessment to a neighbourhood,
ensemble approach. When applying an equivalent-sized neighbourhood to both
configurations, AMM15 typically demonstrated lower (better) scores. These
scores were in turn broken down into off-shelf and on-shelf subdomains and
showed that the different physical processes in these areas affected the
results. Forecast verification studies tailored<?pagebreak page842?> for the coastal/shelf areas
are needed to properly understand the forecast skills in areas with high
complexity and fast-evolving dynamics.</p>
      <p id="d1e1706">When constructing HiRA neighbourhoods, the spatial scales that are
appropriate for the parameter must be considered carefully. This often means
running at several neighbourhood sizes and determining where the scores no
longer seem physically representative. When comparing models, care should be
taken to construct neighbourhood sizes that are similarly sized spatially;
the details of the neighbourhood sizes will depend on the structure and
resolution of the model grid.</p>
      <p id="d1e1710">Treatment of observations is also important in any verification setup. For
this study, the fact that there are different numbers of observations
present at each neighbourhood scale (as observations are rejected due to
land contamination) means that there is never an optimally equalised
dataset (i.e. the same observations for all models and for all neighbourhood
sizes). It also means that comparison of the different neighbourhood results
from a single model is ill advised, in this case, as the observations
numbers can be very different, and therefore the model forecast is being
sampled at different locations. Despite this, observation numbers should be
similar when looking at matched spatially sized neighbourhoods from
different models if results are to be compared. One of the main constraints
identified through this work is both the sparsity and geographical
distribution of observations throughout the NWS domain, with
several viable locations rejected during the HiRA processing due to their
proximity to coastlines.</p>
      <p id="d1e1713">The purest assessment, in terms of observations, would involve a fixed set
of observations, equalised across both model<?pagebreak page843?> configurations and all
neighbourhoods at every time. This would remove the variation in observation
numbers seen as neighbourhood sizes increase as well as those seen between
the two models and give a clean comparison between two models.</p>
      <p id="d1e1716">Care should be taken when applying strict equalisation rules, as this could
result in only a small number of observations being used. The total number
of observations used should be large enough to ensure that the sample is
large enough to produce robust results and satisfy rules for statistical
significance. Equalisation rules could also unfairly affect the spatial
sampling of the verification domain. For example, in this study, coastal
observations would be affected more than deep ocean observations if
neighbourhood equalisation were applied, due to the proximity of the coast.</p>
      <p id="d1e1719">To a lesser extent, the variation in observation numbers on a day-to-day
timescale also has an impact on any results and could mean that incorrect
importance is attributed to certain results, which are simply due to
fluctuations in observation numbers.</p>
      <p id="d1e1722">The fact that the errors can be reduced through the use of neighbourhoods
shows that the ocean and the atmosphere have similarities in the way the
forecasts behave as a function of resolution. This study did not consider
the concept of skill, which incorporates the performance of the forecast
relative to a pre-defined benchmark. For the ocean, the choice of reference
needs to be considered. This could be the subject of further work.</p>
      <p id="d1e1725">To our knowledge, this work is the first attempt to use neighbourhood
techniques to assess ocean models. The promising results showing reductions
in errors of the finer-resolution configuration warrant further work. We see
a number of directions the current study could be extended.</p>
      <p id="d1e1729">The study was conducted on daily output which should be appropriate to
address eddy mesoscale variability, but observations are distributed at
hourly resolution, and so the next logical step would be to assess the
hourly forecasts against the hourly observation and see how this impacted
the results. This will increase the sample size, if all hourly observations
were considered together. However, it is impossible to speculate on whether
considering hourly forecasts would lead to more noisy statistics,
counteracting the larger sample size.</p>
      <p id="d1e1732">This assessment only looked at SST for this initial examination.
Consideration of other ocean variables would also be of interest, including
looking at derived diagnostics such as mixed layer depth, but the sparsity
of observations available for some variables may limit the case studies
available. HiRA as a framework is not remaining static. Enhancements to
introduce non-regular flow-dependent neighbourhoods are planned and may be
of benefit to ocean applications in the future. Finally, an advantage of
using the HiRA framework is that results obtained from deterministic ocean
models could also be compared against results from ensemble models when
these become available for ocean applications.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e1739">Data used in this paper was downloaded from the Copernicus Marine and Environment Monitoring Service (CMEMS).</p>

      <p id="d1e1742">The datasets used were <uri>https://resources.marine.copernicus.eu/?option=com_csw&amp;task=results?option=com_csw&amp;view=details&amp;product_id=NORTHWESTSHELF_ANALYSIS_FORECAST_PHYS_004_001_b</uri> (last access: October 2019),  <uri>https://resources.marine.copernicus.eu/?option=com_csw&amp;task=results?option=com_csw&amp;view=details&amp;product_id=NORTHWESTSHELF_ANALYSIS_FORECAST_PHY_004_013</uri> (last access: October 2019)
and <uri>https://resources.marine.copernicus.eu/?option=com_csw&amp;task=results?option=com_csw&amp;view=details&amp;product_id=INSITU_NWS_NRT_OBSERVATIONS_013_036</uri> (last access: October 2019).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e1757">All authors contributed to the introduction, data and methods, and
conclusions. RC, JM and MM contributed to the scientific evaluation and
analysis of the results. RC and JM designed and ran the model assessments.
CP supported the assessments through the provision and reformatting of the
data used. MT provided details on the model configurations used.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e1763">The authors declare that they have no conflict of interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1769">This study has been conducted using EU Copernicus Marine Service
Information.</p><p id="d1e1771">This work has been carried out as part of the Copernicus Marine Environment
Monitoring Service (CMEMS) HiVE project. CMEMS is implemented by
Mercator Ocean International in the framework of a delegation agreement with
the European Union.</p><p id="d1e1773">Model Evaluation Tools (MET) was developed at the National Center for
Atmospheric Research (NCAR) through grants from the National Science
Foundation (NSF), the National Oceanic and Atmospheric Administration
(NOAA), the United States Air Force (USAF) and the United States Department
of Energy (DOE). NCAR is sponsored by the United States National Science
Foundation.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e1778">This paper was edited by Anna Rubio and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><?label 1?><mixed-citation>Aznar, R., Sotillo, M., Cailleau, S., Lorente, P., Levier, B.,
Amo-Baladrón, A., Reffray, G., and Alvarez Fanjul, E.: Strengths and
weaknesses of the CMEMS forecasted and reanalyzed solutions for the
Iberia-Biscay-Ireland (IBI) waters, J. Marine. Syst., 159, 1–14, <ext-link xlink:href="https://doi.org/10.1016/j.jmarsys.2016.02.007" ext-link-type="DOI">10.1016/j.jmarsys.2016.02.007</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><?label 1?><mixed-citation>Brassington, G.: Forecast Errors, Goodness, and Verification in Ocean
Forecasting, J. Marine Res., 75, 403–433, <ext-link xlink:href="https://doi.org/10.1357/002224017821836851" ext-link-type="DOI">10.1357/002224017821836851</ext-link>, 2017.</mixed-citation></ref>
      <?pagebreak page844?><ref id="bib1.bib3"><label>3</label><?label 1?><mixed-citation>Brier, G. W.: Verification of Forecasts Expressed in Terms of Probability,
Mon. Weather Rev., 78, 1–3, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(1950)078&lt;0001:VOFEIT&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(1950)078&lt;0001:VOFEIT&gt;2.0.CO;2</ext-link>, 1950.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><?label 1?><mixed-citation>Brown, T. A.: Admissible scoring systems for continuous distributions, Santa
Monica, CA, RAND Corporation, available at:
<uri>https://www.rand.org/pubs/papers/P5235.html</uri> (last access: March 2020), 1974.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><?label 1?><mixed-citation>Casati, B., Ross, G., and Stephenson, D. B.: A new intensity-scale approach
for the verification of spatial precipitation forecasts, Met. Apps., 11,
141–154, <ext-link xlink:href="https://doi.org/10.1017/S1350482704001239" ext-link-type="DOI">10.1017/S1350482704001239</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><?label 1?><mixed-citation>Davis, C., Brown, B., and Bullock, R.: Object-Based Verification of
Precipitation Forecasts. Part I: Methodology and Application to Mesoscale
Rain Areas, Mon. Weather Rev., 134, 1772–1784, <ext-link xlink:href="https://doi.org/10.1175/MWR3145.1" ext-link-type="DOI">10.1175/MWR3145.1</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><?label 1?><mixed-citation>Dorninger, M., Gilleland, E., Casati, B., Mittermaier, M. P., Ebert, E. E.,
Brown, B. G., and Wilson, L. J.: The Setup of the MesoVICT Project, B.
Am. Meteorol. Soc., 99, 1887–1906, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-17-0164.1" ext-link-type="DOI">10.1175/BAMS-D-17-0164.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><?label 1?><mixed-citation>Ebert, E. E.: Fuzzy verification of high-resolution gridded forecasts: a
review and proposed framework, Met. Apps, 15, 51–64, <ext-link xlink:href="https://doi.org/10.1002/met.25" ext-link-type="DOI">10.1002/met.25</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><?label 1?><mixed-citation>
Efron, B. and Tibshirani, R.: Bootstrap methods for standard errors,
confidence intervals, and other measures of statistical accuracy,
Stat. Sci., 1, 54–77, 1986.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><?label 1?><mixed-citation>
Epstein, E. S.: A Scoring System for Probability Forecasts of Ranked
Categories, J. Appl. Meteorol., 8, 985–987, 1969.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><?label 1?><mixed-citation>Gilleland, E., Ahijevych, D., Brown, B. G., Casati, B., and Ebert, E. E.:
Intercomparison of Spatial Forecast Verification Methods, Weather Forecast.,
24, 1416–1430, <ext-link xlink:href="https://doi.org/10.1175/2009WAF2222269.1" ext-link-type="DOI">10.1175/2009WAF2222269.1</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><?label 1?><mixed-citation>Goessling, H. F., Tietsche, S., Day, J. J., Hawkins, E., and Jung, T. :
Predictability of the Arctic sea ice edge, Geophys. Res. Lett., 43, 1642–1650, <ext-link xlink:href="https://doi.org/10.1002/2015GL067232" ext-link-type="DOI">10.1002/2015GL067232</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><?label 1?><mixed-citation>Graham, J. A., O'Dea, E., Holt, J., Polton, J., Hewitt, H. T., Furner, R., Guihou, K., Brereton, A., Arnold, A., Wakelin, S., Castillo Sanchez, J. M., and Mayorga Adame, C. G.: AMM15: a new high-resolution NEMO configuration for operational simulation of the European north-west shelf, Geosci. Model Dev., 11, 681–696, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-681-2018" ext-link-type="DOI">10.5194/gmd-11-681-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><?label 1?><mixed-citation>Hernandez, F., Blockley, E., Brassington, G. B., Davidson, F., Divakaran,
P., Drévillon, M., Ishizaki, S., Garcia-Sotillo, M., Hogan, P. J.,
Lagemaa, P., Levier, B., Martin, M., Mehra, A., Mooers, C., Ferry, N., Ryan,
A., Regnier, C., Sellar, A., Smith, G. C., Sofianos, S., Spindler, T.,
Volpe, G., Wilkin, J., Zaron, E. D., and Zhang, A.: Recent progress in
performance evaluations and near real-time assessment of operational ocean
products, J. Oper. Oceanogr., 8, 221–238, <ext-link xlink:href="https://doi.org/10.1080/1755876X.2015.1050282" ext-link-type="DOI">10.1080/1755876X.2015.1050282</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><?label 1?><mixed-citation>Hernandez, F., Smith, G., Baetens, K., Cossarini, G., Garcia-Hermosa, I.,
Drevillon, M., Maksymczuk, J., Melet, A., Regnier, C., and von Schuckmann,
K.: Measuring Performances, Skill and Accuracy in Operational Oceanography:
New Challenges and Approaches, in: New Frontiers in Operational
Oceanography, edited by: Chassignet, E., Pascual, A., Tintore, J., and Verron, J.,
GODAE OceanView, 759–796, <ext-link xlink:href="https://doi.org/10.17125/gov2018.ch29" ext-link-type="DOI">10.17125/gov2018.ch29</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><?label 1?><mixed-citation>Hersbach, H.: Decomposition of the Continuous Ranked Probability Score for
Ensemble Prediction Systems, Weather Forecast., 15, 559–570, <ext-link xlink:href="https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><?label 1?><mixed-citation>Holt, J., Hyder, P., Ashworth, M., Harle, J., Hewitt, H. T., Liu, H., New, A. L., Pickles, S., Porter, A., Popova, E., Allen, J. I., Siddorn, J., and Wood, R.: Prospects for improving the representation of coastal and shelf seas in global ocean models, Geosci. Model Dev., 10, 499–523, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-499-2017" ext-link-type="DOI">10.5194/gmd-10-499-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><?label 1?><mixed-citation>Howarth, M. and Pugh, D.: Chapter 4 Observations of Tides Over the
Continental Shelf of North-West Europe, Elsevier Oceanography Series, 35,
135–188, <ext-link xlink:href="https://doi.org/10.1016/S0422-9894(08)70502-6" ext-link-type="DOI">10.1016/S0422-9894(08)70502-6</ext-link>, 1983.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><?label 1?><mixed-citation>Juza, M., Mourre, B., Lellouche, J. M., Tonani M., and Tintoré, J.: From
basin to sub-basin scale assessment and intercomparison of numerical
simulations in the western Mediterranean Sea, J. Mar. Syst., 149, 36–49, <ext-link xlink:href="https://doi.org/10.1016/j.jmarsys.2015.04.010" ext-link-type="DOI">10.1016/j.jmarsys.2015.04.010</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><?label 1?><mixed-citation>Keil, C. and Craig, G. C.: A Displacement-Based Error Measure Applied in a
Regional Ensemble Forecasting System, Mon. Weather Rev., 135, 3248–3259,
<ext-link xlink:href="https://doi.org/10.1175/MWR3457.1" ext-link-type="DOI">10.1175/MWR3457.1</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><?label 1?><mixed-citation>
King, R., While, J., Martin, M. J., Lea, D. J., Lemieux-Dudon, B., Waters,
J., and O'Dea, E.: Improving the initialisation of the Met Office operational
shelf-seas model, Ocean Model., 130, 1–14, 2018.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><?label 1?><mixed-citation>Le Traon, P. Y., Reppucci, A., Alvarez Fanjul, E., Aouf, L., Behrens, A., Belmonte, M.,
Bentamy, A., Bertino, L., Brando, V. E., Kreiner, M. B., Benkiran, M., Carval, T., Ciliberti,
S. A., Claustre, H., Clementi, E., Coppini, G., Cossarini, G., De Alfonso
Alonso-Muñoyerro, M., Delamarche, A., Dibarboure, G., Dinessen, F., Drevillon, M.,
Drillet, Y., Faugere, Y., Fernández, V., Fleming, A., Garcia-Hermosa, M. I., Sotillo,
M. G., Garric, G., Gasparin, F., Giordan, C., Gehlen, M., Gregoire, M. L., Guinehut, S.,
Hamon, M., Harris, C., Hernandez, F., Hinkler, J. B., Hoyer, J., Karvonen, J., Kay, S., King,
R., Lavergne, T., Lemieux-Dudon, B., Lima, L., Mao, C., Martin, M. J., Masina, S., Melet, A.,
Buongiorno Nardelli, B., Nolan, G., Pascual, A., Pistoia, J., Palazov, A., Piolle, J. F.,
Pujol, M. I., Pequignet, A. C., Peneva, E., Pérez Gómez, B., Petit de la Villeon,
L., Pinardi, N., Pisano, A., Pouliquen, S., Reid, R., Remy, E., Santoleri, R., Siddorn, J.,
She, J., Staneva, J., Stoffelen, A., Tonani, M., Vandenbulcke, L., von Schuckmann, K.,
Volpe, G., Wettre, C., and Zacharioudaki, A.: From Observation to Information and
Users: The Copernicus Marine Service Perspective, Front. Mar. Sci., 6, 234, <ext-link xlink:href="https://doi.org/10.3389/fmars.2019.00234" ext-link-type="DOI">10.3389/fmars.2019.00234</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><?label 1?><mixed-citation>Lorente, P., García-Sotillo, M., Amo-Baladrón, A., Aznar, R., Levier, B., Sánchez-Garrido, J. C., Sammartino, S., de Pascual-Collar, Á., Reffray, G., Toledano, C., and Álvarez-Fanjul, E.: Skill assessment of global, regional, and coastal circulation forecast models: evaluating the benefits of dynamical downscaling in IBI (Iberia–Biscay–Ireland) surface waters, Ocean Sci., 15, 967–996, <ext-link xlink:href="https://doi.org/10.5194/os-15-967-2019" ext-link-type="DOI">10.5194/os-15-967-2019</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><?label 1?><mixed-citation>
Lorente, P., Sotillo, M., Amo-Baladrón, A., Aznar, R., Levier, B., Aouf,
L., Dabrowski, T., Pascual, Á., Reffray, G., Dalphinet, A., Toledano
Lozano, C., Rainaud, R., and Alvarez Fanjul, E. : The NARVAL Software
Toolbox in Support of Ocean Models Skill Assessment at Regional and Coastal
Scales, doi10.1007/978-3-030-22747-0_25, 2019b.</mixed-citation></ref>
      <?pagebreak page845?><ref id="bib1.bib25"><label>25</label><?label 1?><mixed-citation>
Madec, G. and the NEMO team: NEMO ocean engine, Note du Pôle de
modélisation, Institut Pierre-Simon Laplace (IPSL), France, No 27 ISSN
1288-1619, 2016.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><?label 1?><mixed-citation>Mason, E., Ruiz, S., Bourdalle-Badie, R., Reffray, G., García-Sotillo, M., and Pascual, A.: New insight into 3-D mesoscale eddy properties from CMEMS operational models in the western Mediterranean, Ocean Sci., 15, 1111–1131, <ext-link xlink:href="https://doi.org/10.5194/os-15-1111-2019" ext-link-type="DOI">10.5194/os-15-1111-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><?label 1?><mixed-citation>Mass, C. F., Ovens, D., Westrick, K., and Colle, B. A.: DOES INCREASING
HORIZONTAL RESOLUTION PRODUCE MORE SKILLFUL FORECASTS?, B. Am. Meteorol.
Soc., 83, 407–430, <ext-link xlink:href="https://doi.org/10.1175/1520-0477(2002)083&lt;0407:DIHRPM&gt;2.3.CO;2" ext-link-type="DOI">10.1175/1520-0477(2002)083&lt;0407:DIHRPM&gt;2.3.CO;2</ext-link>, 2002.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><?label 1?><mixed-citation>Mirouze, I., Blockley, E. W., Lea, D. J., Martin, M. J., and Bell, M. J.: A
multiple length scale correlation operator for ocean data assimilation,
Tellus A, 68, 29744, <ext-link xlink:href="https://doi.org/10.3402/tellusa.v68.29744" ext-link-type="DOI">10.3402/tellusa.v68.29744</ext-link>,
2016.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><?label 1?><mixed-citation>Mittermaier, M., Roberts, N., and Thompson, S. A.: A long-term assessment of
precipitation forecast skill using the Fractions Skill Score, Met. Apps, 20,
176–186, <ext-link xlink:href="https://doi.org/10.1002/met.296" ext-link-type="DOI">10.1002/met.296</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><?label 1?><mixed-citation>Mittermaier, M. P.: A Strategy for Verifying Near-Convection-Resolving Model
Forecasts at Observing Sites, Weather Forecast., 29, 185–204, <ext-link xlink:href="https://doi.org/10.1175/WAF-D-12-00075.1" ext-link-type="DOI">10.1175/WAF-D-12-00075.1</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><?label 1?><mixed-citation>Mittermaier, M. P. and Csima, G.: Ensemble versus Deterministic Performance
at the Kilometer Scale, Weather Forecast., 32, 1697–1709, <ext-link xlink:href="https://doi.org/10.1175/WAF-D-16-0164.1" ext-link-type="DOI">10.1175/WAF-D-16-0164.1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><?label 1?><mixed-citation>
Mogensen, K., Balmaseda, M. A., and Weaver, A.: The NEMOVAR ocean data
assimilation system as implemented in the ECMWF ocean analysis for System 4,
European Centre for Medium-Range Weather Forecasts, Reading, UK, 2012.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><?label 1?><mixed-citation>Mourre, B., Aguiar, E., Juza, M., Hernandez-Lasheras, J., Reyes, E., Heslop, E., Escudier,
R., Cutolo, E., Ruiz, S., Mason, E., Pascual, A., and Tintoré, J.:
Assessment of high-resolution regional ocean prediction systems using
muli-platform observations: illustrations in the Western Mediterranean Sea,
in: New Frontiers in Operational Oceanography, edited by: Chassignet, E., Pascual, A., Tintoré,
J., and Verron, J., GODAE Ocean View, 663–694, <ext-link xlink:href="https://doi.org/10.17125/gov2018.ch24" ext-link-type="DOI">10.17125/gov2018.ch24</ext-link>, 2018.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bib34"><label>34</label><?label 1?><mixed-citation>O'Dea, E. J., Arnold, A. K., Edwards, K. P., Furner, R., Hyder, P., Martin,
M. J., Siddorn, J. R., Storkey, D., While, J., Holt, J. T., and Liu, H.: An
operational ocean forecast system incorporating NEMO and SST data
assimilation for the tidally driven European North-West shelf, J. Oper.
Oceanogr., 5, 3–17, <ext-link xlink:href="https://doi.org/10.1080/1755876X.2012.11020128" ext-link-type="DOI">10.1080/1755876X.2012.11020128</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><?label 1?><mixed-citation>O'Dea, E., Furner, R., Wakelin, S., Siddorn, J., While, J., Sykes, P., King, R., Holt, J., and Hewitt, H.: The CO5 configuration of the 7 km Atlantic Margin Model: large-scale biases and sensitivity to forcing, physics options and vertical resolution, Geosci. Model Dev., 10, 2947–2969, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-2947-2017" ext-link-type="DOI">10.5194/gmd-10-2947-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><?label 1?><mixed-citation>Rossa, A., Nurmi, P., and Ebert, E.: Overview of methods for the verification of
quantitative precipitation forecasts, in: Precipitation: Advances in
Measurement, Estimation and Prediction, edited by: Michaelides, S.,
Springer, Berlin, Heidelberg, 419–452, <ext-link xlink:href="https://doi.org/10.1007/978-3-540-77655-0_16" ext-link-type="DOI">10.1007/978-3-540-77655-0_16</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><?label 1?><mixed-citation>Tonani, M., Sykes, P., King, R. R., McConnell, N., Péquignet, A.-C., O'Dea, E., Graham, J. A., Polton, J., and Siddorn, J.: The impact of a new high-resolution ocean model on the Met Office North-West European Shelf forecasting system, Ocean Sci., 15, 1133–1158, <ext-link xlink:href="https://doi.org/10.5194/os-15-1133-2019" ext-link-type="DOI">10.5194/os-15-1133-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><?label 1?><mixed-citation>World Meteorological Organisation: Guide to Meteorological Instruments and
Methods of Observation (WMO-No. 8, the CIMO Guide), available at: <uri>https://library.wmo.int/opac/doc_num.php?explnum_id=4147</uri> (last access: June 2019), 2017.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>An approach to the verification of high-resolution ocean models using spatial methods</article-title-html>
<abstract-html><p>The Met Office currently runs two operational ocean forecasting
configurations for the  North West European Shelf: an eddy-permitting model
with a resolution of 7&thinsp;km (AMM7) and an eddy-resolving model at 1.5&thinsp;km
(AMM15).</p><p>Whilst qualitative assessments have demonstrated the benefits brought by the
increased resolution of AMM15, particularly in the ability to resolve
finer-scale features, it has been difficult to show this quantitatively,
especially in forecast mode. Applications of typical assessment metrics such
as the root mean square error have been inconclusive, as the high-resolution
model tends to be penalised more severely, referred to as the double-penalty
effect. This effect occurs in point-to-point comparisons whereby features
correctly forecast but misplaced with respect to the observations are
penalised twice: once for not occurring at the observed location, and
secondly for occurring at the forecast location, where they have not been
observed.</p><p>An exploratory assessment of sea surface temperature (SST) has been made at
in situ observation locations using a
single-observation neighbourhood-forecast (SO-NF) spatial verification
method known as the High-Resolution Assessment (HiRA) framework. The primary
focus of the assessment was to capture important aspects of methodology to
consider when applying the HiRA framework. Forecast grid points within
neighbourhoods centred on the observing location are considered as pseudo
ensemble members, so that typical ensemble and probabilistic forecast
verification metrics such as the continuous ranked probability score (CRPS)
can be utilised. It is found that through the application of HiRA it is
possible to identify improvements in the higher-resolution model which were
not apparent using typical grid-scale assessments.</p><p>This work suggests that future comparative assessments of ocean models with
different resolutions would benefit from using HiRA as part of the
evaluation process, as it gives a more equitable and appropriate reflection
of model performance at higher resolutions.</p></abstract-html>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
Aznar, R., Sotillo, M., Cailleau, S., Lorente, P., Levier, B.,
Amo-Baladrón, A., Reffray, G., and Alvarez Fanjul, E.: Strengths and
weaknesses of the CMEMS forecasted and reanalyzed solutions for the
Iberia-Biscay-Ireland (IBI) waters, J. Marine. Syst., 159, 1–14, <a href="https://doi.org/10.1016/j.jmarsys.2016.02.007" target="_blank">https://doi.org/10.1016/j.jmarsys.2016.02.007</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
Brassington, G.: Forecast Errors, Goodness, and Verification in Ocean
Forecasting, J. Marine Res., 75, 403–433, <a href="https://doi.org/10.1357/002224017821836851" target="_blank">https://doi.org/10.1357/002224017821836851</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
Brier, G. W.: Verification of Forecasts Expressed in Terms of Probability,
Mon. Weather Rev., 78, 1–3, <a href="https://doi.org/10.1175/1520-0493(1950)078&lt;0001:VOFEIT&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(1950)078&lt;0001:VOFEIT&gt;2.0.CO;2</a>, 1950.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
Brown, T. A.: Admissible scoring systems for continuous distributions, Santa
Monica, CA, RAND Corporation, available at:
<a href="https://www.rand.org/pubs/papers/P5235.html" target="_blank"/> (last access: March 2020), 1974.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
Casati, B., Ross, G., and Stephenson, D. B.: A new intensity-scale approach
for the verification of spatial precipitation forecasts, Met. Apps., 11,
141–154, <a href="https://doi.org/10.1017/S1350482704001239" target="_blank">https://doi.org/10.1017/S1350482704001239</a>, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
Davis, C., Brown, B., and Bullock, R.: Object-Based Verification of
Precipitation Forecasts. Part I: Methodology and Application to Mesoscale
Rain Areas, Mon. Weather Rev., 134, 1772–1784, <a href="https://doi.org/10.1175/MWR3145.1" target="_blank">https://doi.org/10.1175/MWR3145.1</a>, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
Dorninger, M., Gilleland, E., Casati, B., Mittermaier, M. P., Ebert, E. E.,
Brown, B. G., and Wilson, L. J.: The Setup of the MesoVICT Project, B.
Am. Meteorol. Soc., 99, 1887–1906, <a href="https://doi.org/10.1175/BAMS-D-17-0164.1" target="_blank">https://doi.org/10.1175/BAMS-D-17-0164.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
Ebert, E. E.: Fuzzy verification of high-resolution gridded forecasts: a
review and proposed framework, Met. Apps, 15, 51–64, <a href="https://doi.org/10.1002/met.25" target="_blank">https://doi.org/10.1002/met.25</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
Efron, B. and Tibshirani, R.: Bootstrap methods for standard errors,
confidence intervals, and other measures of statistical accuracy,
Stat. Sci., 1, 54–77, 1986.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
Epstein, E. S.: A Scoring System for Probability Forecasts of Ranked
Categories, J. Appl. Meteorol., 8, 985–987, 1969.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
Gilleland, E., Ahijevych, D., Brown, B. G., Casati, B., and Ebert, E. E.:
Intercomparison of Spatial Forecast Verification Methods, Weather Forecast.,
24, 1416–1430, <a href="https://doi.org/10.1175/2009WAF2222269.1" target="_blank">https://doi.org/10.1175/2009WAF2222269.1</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
Goessling, H. F., Tietsche, S., Day, J. J., Hawkins, E., and Jung, T. :
Predictability of the Arctic sea ice edge, Geophys. Res. Lett., 43, 1642–1650, <a href="https://doi.org/10.1002/2015GL067232" target="_blank">https://doi.org/10.1002/2015GL067232</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
Graham, J. A., O'Dea, E., Holt, J., Polton, J., Hewitt, H. T., Furner, R., Guihou, K., Brereton, A., Arnold, A., Wakelin, S., Castillo Sanchez, J. M., and Mayorga Adame, C. G.: AMM15: a new high-resolution NEMO configuration for operational simulation of the European north-west shelf, Geosci. Model Dev., 11, 681–696, <a href="https://doi.org/10.5194/gmd-11-681-2018" target="_blank">https://doi.org/10.5194/gmd-11-681-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
Hernandez, F., Blockley, E., Brassington, G. B., Davidson, F., Divakaran,
P., Drévillon, M., Ishizaki, S., Garcia-Sotillo, M., Hogan, P. J.,
Lagemaa, P., Levier, B., Martin, M., Mehra, A., Mooers, C., Ferry, N., Ryan,
A., Regnier, C., Sellar, A., Smith, G. C., Sofianos, S., Spindler, T.,
Volpe, G., Wilkin, J., Zaron, E. D., and Zhang, A.: Recent progress in
performance evaluations and near real-time assessment of operational ocean
products, J. Oper. Oceanogr., 8, 221–238, <a href="https://doi.org/10.1080/1755876X.2015.1050282" target="_blank">https://doi.org/10.1080/1755876X.2015.1050282</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
Hernandez, F., Smith, G., Baetens, K., Cossarini, G., Garcia-Hermosa, I.,
Drevillon, M., Maksymczuk, J., Melet, A., Regnier, C., and von Schuckmann,
K.: Measuring Performances, Skill and Accuracy in Operational Oceanography:
New Challenges and Approaches, in: New Frontiers in Operational
Oceanography, edited by: Chassignet, E., Pascual, A., Tintore, J., and Verron, J.,
GODAE OceanView, 759–796, <a href="https://doi.org/10.17125/gov2018.ch29" target="_blank">https://doi.org/10.17125/gov2018.ch29</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
Hersbach, H.: Decomposition of the Continuous Ranked Probability Score for
Ensemble Prediction Systems, Weather Forecast., 15, 559–570, <a href="https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0434(2000)015&lt;0559:DOTCRP&gt;2.0.CO;2</a>, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
Holt, J., Hyder, P., Ashworth, M., Harle, J., Hewitt, H. T., Liu, H., New, A. L., Pickles, S., Porter, A., Popova, E., Allen, J. I., Siddorn, J., and Wood, R.: Prospects for improving the representation of coastal and shelf seas in global ocean models, Geosci. Model Dev., 10, 499–523, <a href="https://doi.org/10.5194/gmd-10-499-2017" target="_blank">https://doi.org/10.5194/gmd-10-499-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
Howarth, M. and Pugh, D.: Chapter 4 Observations of Tides Over the
Continental Shelf of North-West Europe, Elsevier Oceanography Series, 35,
135–188, <a href="https://doi.org/10.1016/S0422-9894(08)70502-6" target="_blank">https://doi.org/10.1016/S0422-9894(08)70502-6</a>, 1983.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
Juza, M., Mourre, B., Lellouche, J. M., Tonani M., and Tintoré, J.: From
basin to sub-basin scale assessment and intercomparison of numerical
simulations in the western Mediterranean Sea, J. Mar. Syst., 149, 36–49, <a href="https://doi.org/10.1016/j.jmarsys.2015.04.010" target="_blank">https://doi.org/10.1016/j.jmarsys.2015.04.010</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
Keil, C. and Craig, G. C.: A Displacement-Based Error Measure Applied in a
Regional Ensemble Forecasting System, Mon. Weather Rev., 135, 3248–3259,
<a href="https://doi.org/10.1175/MWR3457.1" target="_blank">https://doi.org/10.1175/MWR3457.1</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
King, R., While, J., Martin, M. J., Lea, D. J., Lemieux-Dudon, B., Waters,
J., and O'Dea, E.: Improving the initialisation of the Met Office operational
shelf-seas model, Ocean Model., 130, 1–14, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
Le Traon, P. Y., Reppucci, A., Alvarez Fanjul, E., Aouf, L., Behrens, A., Belmonte, M.,
Bentamy, A., Bertino, L., Brando, V. E., Kreiner, M. B., Benkiran, M., Carval, T., Ciliberti,
S. A., Claustre, H., Clementi, E., Coppini, G., Cossarini, G., De Alfonso
Alonso-Muñoyerro, M., Delamarche, A., Dibarboure, G., Dinessen, F., Drevillon, M.,
Drillet, Y., Faugere, Y., Fernández, V., Fleming, A., Garcia-Hermosa, M. I., Sotillo,
M. G., Garric, G., Gasparin, F., Giordan, C., Gehlen, M., Gregoire, M. L., Guinehut, S.,
Hamon, M., Harris, C., Hernandez, F., Hinkler, J. B., Hoyer, J., Karvonen, J., Kay, S., King,
R., Lavergne, T., Lemieux-Dudon, B., Lima, L., Mao, C., Martin, M. J., Masina, S., Melet, A.,
Buongiorno Nardelli, B., Nolan, G., Pascual, A., Pistoia, J., Palazov, A., Piolle, J. F.,
Pujol, M. I., Pequignet, A. C., Peneva, E., Pérez Gómez, B., Petit de la Villeon,
L., Pinardi, N., Pisano, A., Pouliquen, S., Reid, R., Remy, E., Santoleri, R., Siddorn, J.,
She, J., Staneva, J., Stoffelen, A., Tonani, M., Vandenbulcke, L., von Schuckmann, K.,
Volpe, G., Wettre, C., and Zacharioudaki, A.: From Observation to Information and
Users: The Copernicus Marine Service Perspective, Front. Mar. Sci., 6, 234, <a href="https://doi.org/10.3389/fmars.2019.00234" target="_blank">https://doi.org/10.3389/fmars.2019.00234</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
Lorente, P., García-Sotillo, M., Amo-Baladrón, A., Aznar, R., Levier, B., Sánchez-Garrido, J. C., Sammartino, S., de Pascual-Collar, Á., Reffray, G., Toledano, C., and Álvarez-Fanjul, E.: Skill assessment of global, regional, and coastal circulation forecast models: evaluating the benefits of dynamical downscaling in IBI (Iberia–Biscay–Ireland) surface waters, Ocean Sci., 15, 967–996, <a href="https://doi.org/10.5194/os-15-967-2019" target="_blank">https://doi.org/10.5194/os-15-967-2019</a>, 2019a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
Lorente, P., Sotillo, M., Amo-Baladrón, A., Aznar, R., Levier, B., Aouf,
L., Dabrowski, T., Pascual, Á., Reffray, G., Dalphinet, A., Toledano
Lozano, C., Rainaud, R., and Alvarez Fanjul, E. : The NARVAL Software
Toolbox in Support of Ocean Models Skill Assessment at Regional and Coastal
Scales, doi10.1007/978-3-030-22747-0_25, 2019b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
Madec, G. and the NEMO team: NEMO ocean engine, Note du Pôle de
modélisation, Institut Pierre-Simon Laplace (IPSL), France, No 27 ISSN
1288-1619, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
Mason, E., Ruiz, S., Bourdalle-Badie, R., Reffray, G., García-Sotillo, M., and Pascual, A.: New insight into 3-D mesoscale eddy properties from CMEMS operational models in the western Mediterranean, Ocean Sci., 15, 1111–1131, <a href="https://doi.org/10.5194/os-15-1111-2019" target="_blank">https://doi.org/10.5194/os-15-1111-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
Mass, C. F., Ovens, D., Westrick, K., and Colle, B. A.: DOES INCREASING
HORIZONTAL RESOLUTION PRODUCE MORE SKILLFUL FORECASTS?, B. Am. Meteorol.
Soc., 83, 407–430, <a href="https://doi.org/10.1175/1520-0477(2002)083&lt;0407:DIHRPM&gt;2.3.CO;2" target="_blank">https://doi.org/10.1175/1520-0477(2002)083&lt;0407:DIHRPM&gt;2.3.CO;2</a>, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
Mirouze, I., Blockley, E. W., Lea, D. J., Martin, M. J., and Bell, M. J.: A
multiple length scale correlation operator for ocean data assimilation,
Tellus A, 68, 29744, <a href="https://doi.org/10.3402/tellusa.v68.29744" target="_blank">https://doi.org/10.3402/tellusa.v68.29744</a>,
2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
Mittermaier, M., Roberts, N., and Thompson, S. A.: A long-term assessment of
precipitation forecast skill using the Fractions Skill Score, Met. Apps, 20,
176–186, <a href="https://doi.org/10.1002/met.296" target="_blank">https://doi.org/10.1002/met.296</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
Mittermaier, M. P.: A Strategy for Verifying Near-Convection-Resolving Model
Forecasts at Observing Sites, Weather Forecast., 29, 185–204, <a href="https://doi.org/10.1175/WAF-D-12-00075.1" target="_blank">https://doi.org/10.1175/WAF-D-12-00075.1</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
Mittermaier, M. P. and Csima, G.: Ensemble versus Deterministic Performance
at the Kilometer Scale, Weather Forecast., 32, 1697–1709, <a href="https://doi.org/10.1175/WAF-D-16-0164.1" target="_blank">https://doi.org/10.1175/WAF-D-16-0164.1</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
Mogensen, K., Balmaseda, M. A., and Weaver, A.: The NEMOVAR ocean data
assimilation system as implemented in the ECMWF ocean analysis for System 4,
European Centre for Medium-Range Weather Forecasts, Reading, UK, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
Mourre, B., Aguiar, E., Juza, M., Hernandez-Lasheras, J., Reyes, E., Heslop, E., Escudier,
R., Cutolo, E., Ruiz, S., Mason, E., Pascual, A., and Tintoré, J.:
Assessment of high-resolution regional ocean prediction systems using
muli-platform observations: illustrations in the Western Mediterranean Sea,
in: New Frontiers in Operational Oceanography, edited by: Chassignet, E., Pascual, A., Tintoré,
J., and Verron, J., GODAE Ocean View, 663–694, <a href="https://doi.org/10.17125/gov2018.ch24" target="_blank">https://doi.org/10.17125/gov2018.ch24</a>, 2018.

</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
O'Dea, E. J., Arnold, A. K., Edwards, K. P., Furner, R., Hyder, P., Martin,
M. J., Siddorn, J. R., Storkey, D., While, J., Holt, J. T., and Liu, H.: An
operational ocean forecast system incorporating NEMO and SST data
assimilation for the tidally driven European North-West shelf, J. Oper.
Oceanogr., 5, 3–17, <a href="https://doi.org/10.1080/1755876X.2012.11020128" target="_blank">https://doi.org/10.1080/1755876X.2012.11020128</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
O'Dea, E., Furner, R., Wakelin, S., Siddorn, J., While, J., Sykes, P., King, R., Holt, J., and Hewitt, H.: The CO5 configuration of the 7&thinsp;km Atlantic Margin Model: large-scale biases and sensitivity to forcing, physics options and vertical resolution, Geosci. Model Dev., 10, 2947–2969, <a href="https://doi.org/10.5194/gmd-10-2947-2017" target="_blank">https://doi.org/10.5194/gmd-10-2947-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
Rossa, A., Nurmi, P., and Ebert, E.: Overview of methods for the verification of
quantitative precipitation forecasts, in: Precipitation: Advances in
Measurement, Estimation and Prediction, edited by: Michaelides, S.,
Springer, Berlin, Heidelberg, 419–452, <a href="https://doi.org/10.1007/978-3-540-77655-0_16" target="_blank">https://doi.org/10.1007/978-3-540-77655-0_16</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
Tonani, M., Sykes, P., King, R. R., McConnell, N., Péquignet, A.-C., O'Dea, E., Graham, J. A., Polton, J., and Siddorn, J.: The impact of a new high-resolution ocean model on the Met Office North-West European Shelf forecasting system, Ocean Sci., 15, 1133–1158, <a href="https://doi.org/10.5194/os-15-1133-2019" target="_blank">https://doi.org/10.5194/os-15-1133-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
World Meteorological Organisation: Guide to Meteorological Instruments and
Methods of Observation (WMO-No. 8, the CIMO Guide), available at: <a href="https://library.wmo.int/opac/doc_num.php?explnum_id=4147" target="_blank"/> (last access: June 2019), 2017.
</mixed-citation></ref-html>--></article>
