Improving usability and pregnancy rates of a fertility monitor by an additional mobile application: results of a retrospective efficacy study of Daysy and DaysyView app
Koch MC, Lermann J, van de Roemer N, Renner SK, Burghaus S, Hackl J, Dittrich R, Kehl S, Oppelt PG, Hildebrandt T, Hack CC, Pöhls UG, Renner SP, Thiel FC.
- DOI
- 10.1186/s12978-018-0479-6
- Record issued
- 2026-08-05
- Engine
- 7.15.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/00d56b27-8f31-470a-84c5-aeaa42fa8266 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×7−3.5★
- ClaimsOverstated claim ×2−1★
- ReportingStatistical analysis not met−0.5★
- ReportingData & code availability not met−0.5★
- CopyeditManuscript needs significant editing−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×2−0.25★
- ReportingStudy design partially met−0.25★
- ReportingKey resources partially met−0.25★
- No data or code availability links were detected to verify.
- 01Statistical reporting inadequate
The paper has a demonstrable error in the confidence intervals in Table 2 (reversed values), and it lacks reporting of test assumptions and software.
“****t-test p = 0.0001”
Figure 1 - 02Data and code not shared
The data availability statement is vague ('contact author for data request') with no mechanism or timeframe, and no code is shared.
“Please contact author for data request.”
Data availability - 03Printed percentage does not match its own count
63.2% does not match the reported count 506/798
- 04Printed percentage does not match its own count
10% does not match the reported count 2/22
“10% (n = 2)”
In women over 40 (n=22), 10% (n=2) indi… - 05Mathematically impossible statistic
In Table 2, the column labelled 'CI, lower Limit' contains values larger than those in 'CI, upper Limit' for every row, which is impossible for a confidence interval.
“| 1 | 696 | 4 | 0.57 | 1.57 | 0.18 |”
Table 2 - 06Internal contradictions in the reported numbers
Confidence intervals in Table 2 are reversed, with lower limits exceeding upper limits, indicating a data presentation error.
CI, lower Limit (%) | CI, upper Limit (%) | ... 1 | 696 | 4 | 0.57 | 1.57 | 0.18
Table 2reviewer’s wording
2 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper has a clear scientific premise and adequate ethical approvals, but contains demonstrable errors in Table 2 and internal inconsistencies in reported numbers, lacks statistical software and data access details, and has gaps in study design reporting. These issues reduce the robustness of the published findings.
Three independent reviewer runs (all from the same model) were synthesized; their agreement on most dimensions supports stability. The statistics component recomputed only 2 tests, of which 0 were consistent; this limited coverage should not be interpreted as an endorsement of the rest. The citation check found no retracted or non-existent references.
Numerical inconsistencies
3 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- Mathematically impossible statisticAssessed
- Printed percentage does not match its own countRecomputed
2 reported summary statistics mathematically impossible for the stated N (PERCENT).
- PERCENT63.2% does not match the reported count 506/798
- PERCENT10% does not match the reported count 2/22
“10% (n = 2)”
In women over 40 (n=22), 10% (n=2) indi…
- mediumimpossible statisticIn Table 2, the column labelled 'CI, lower Limit' contains values larger than those in 'CI, upper Limit' for every row, which is impossible for a confidence interval.
“| 1 | 696 | 4 | 0.57 | 1.57 | 0.18 |”
Table 2 - mediuminternal contradictionConfidence intervals in Table 2 are reversed, with lower limits exceeding upper limits, indicating a data presentation error.
CI, lower Limit (%) | CI, upper Limit (%) | ... 1 | 696 | 4 | 0.57 | 1.57 | 0.18
Table 2reviewer’s wording - lowinternal contradiction524 women indicated additional contraceptive use, but the next sentence analyzes 493 respondents using additional methods; the difference is unexplained.
From 798 participating women, 524 (64%) indicated an additional contraceptive use (Fig. ). Out of the 493 respondents using additional precaution methods, 73% (358 out of 493)...
Results ¶2reviewer’s wording - lowinternal contradictionSeveral percentages do not match their reported counts (e.g., 524/798 = 65.7% not 64%; 239/798 = 29.95% not 29.24%; 69/798 = 8.65% not 9.01%).
From 798 participating women, 524 (64%) indicated an additional contraceptive use ... 239 (29.24%) ... 69 (9.01%)
Resultsreviewer’s wording - lowinternal contradictionThe total number of recorded cycles is reported as 4738 in the first Results paragraph and as 4750 later, a discrepancy of 12 cycles.
The total number of recorded cycles was 4738. ... a total of 4750 cycles were identified.
Results ¶1reviewer’s wording - lowinternal contradictionAccount numbers are inconsistent: 776 ready + 20 deleted = 796, but 798 participants and 778 remaining accounts are stated.
1 year after the study was started (November 1st, 2016) the status of 776 (98%) DaysyView accounts is “Ready” ... 20 (2%) accounts ... Of the 778 remaining accounts, 618 (79%)...
Resultsreviewer’s wording
Overstated conclusions
2 findings · worst mediumConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
12 major claims checked against the paper's own evidence: 2 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewers 1, 2The typical-use related Pearl-Index significantly improved from 3.8 to 1.3.The improvement may be real, but the word 'significantly' is not supported by any statistical test comparing independent cohorts.Evidence: This paper reports a typical-use PI of 1.25 (2 pregnancies/2076 cycles); the 3.8 value is from the historical Freundl study. No statistical test comparing the two cohorts is reported.
“However, if the focus is on the typical-use related Pearl-Index , it has significantly improved from 3,8 to 1,3.”
Abstract - overstatedReviewer 2Combining Daysy with the DaysyView app improves usability and enhances typical-, method-, and perfect-use pregnancy rates.The observed pregnancy rates are consistent with the claim, but the lack of a control group and the reliance on historical comparators do not support a causal attribution of the improvement to the app.Evidence: Single-arm retrospective survey reporting Pearl Index values of 1.3 (typical-use), 0.6 (method-use), and 0.8 (perfect-use); no concurrent control group using Daysy alone.
“We conclude that it is possible through the present technology of Daysy and the additional, optional use of DaysyView, to improve usability and enhance the usage safety as well as the method safety rate.”
Conclusion - partialReviewer 1Combining Daysy with DaysyView results in higher overall usability and improved pregnancy rates.The study shows improved PI compared to historical data, but the retrospective design and lack of a control group limit causal inference; other factors (device improvements, demographic changes) may contribute.Evidence: Typical-use PI improved from 3.8 to 1.3 compared to Freundl et al., and app usage was high (65% daily, 84% better understanding).
“It seems that combining a specific biosensor-embedded device (Daysy), which gives the method a very high repeatable accuracy, and a mobile application (DaysyView) which leads to higher user engagement, results in higher overall usability of the method.”
Discussion - partialReviewer 1It is possible through the present technology of Daysy and the additional, optional use of DaysyView to improve usability and enhance the pregnancy rates.The study provides evidence of improved PI and high app engagement, but the retrospective design and confounders prevent strong causal conclusions.Evidence: Improved PI and high app usage rates.
“We conclude, that it is possible through the present technology of Daysy and the additional, optional use of DaysyView to improve usability and enhance the typical-, method- and perfect -use pregnancy rates.”
Conclusion - partialReviewer 3Through the additional use of an App and thereby improved usability of the medical device, it is possible to enhance the typical-use related as well as the method-related pregnancy rates.The claim is partially supported by comparing the current typical-use Pearl Index (1.3) to a historical control (3.8 from Freundl et al.), but no direct comparison group using the device without the app is included, so the improvement may be due to other factors.Evidence: Typical-use PI of 1.3 vs. historical 3.8; method-related PI of 0.6 vs. 0.7.
“In the resultant group of 125 women (2076 cycles in total), 2 women indicated that they had been unintentionally pregnant during the use of the device, giving a typical-use related Pearl-Index of 1.3.”
Abstract - partialReviewer 3It seems that combining a specific biosensor-embedded device (Daysy), which gives the method a very high repeatable accuracy, and a mobile application (DaysyView) which leads to higher user engagement, results in higher overall usability of the method.The paper shows high app usage and user satisfaction, but does not directly measure usability or compare it to a group without the app. The conclusion is plausible but not definitively proven.Evidence: 65% of participants use the app daily; 84% report better understanding of their cycle.
“It seems that combining a specific biosensor-embedded device (Daysy), which gives the method a very high repeatable accuracy, and a mobile application (DaysyView) which leads to higher user engagement, results in higher overall usability of the method.”
Conclusion - supportedReviewer 1The typical-use Pearl-Index is 1.3.The calculation from the reported data (2 pregnancies, 2076 cycles) gives 1.252, rounded to 1.3, which is consistent.Evidence: 2 pregnancies in 2076 cycles, calculation: 2×1300/2076≈1.25.
“In the resultant group of 125 women (2076 cycles in total), 2 women indicated that they had been unintentionally pregnant during the use of the device, giving a typical-use related Pearl-Index of 1.3.”
Abstract - supportedReviewer 1The method-related Pearl-Index is 0.6.Based on 1 pregnancy during green phase, calculation gives 0.626, rounded to 0.6.Evidence: 1 pregnancy during green phase, 2076 cycles, calculation: 1×1300/2076≈0.63.
“Counting only the pregnancies which occurred as a result of unprotected intercourse during the infertile (green) phase, we found 1 pregnancy, giving a method-related Pearl-Index of 0.6.”
Abstract - supportedReviewer 1The perfect-use Pearl-Index is 0.8.Based on 1 pregnancy in 1725 cycles, calculation gives 0.753, rounded to 0.8.Evidence: 1 pregnancy in 1725 cycles, calculation: 1×1300/1725≈0.75.
“Calculating the pregnancy rate resulting from continuous use and unprotected intercourse exclusively on green days, gives a perfect-use Pearl-Index of 0.8.”
Abstract - supportedReviewer 2The perfect-use efficacy of Daysy is 0.8.The calculation is internally consistent and supports the claim within rounding.Evidence: Results state: 'the perfect-use pregnancy rate is 1 × 1300 / 1725, which equals a PI of 0.753.'
“Independently, the perfect-use efficacy (0,8) of Daysy was calculated in this study.”
Abstract - supportedReviewer 2The method-related Pearl-Index (0.6) differs only a little from Freundl's reported 0.7.The numbers are close and the claim is modest, so the evidence supports it.Evidence: Method-related PI in this study is 0.626 (1 pregnancy/2076 cycles); the cited Freundl value is 0.7.
“The result of the method related Pearl-Index calculation obtained in the present study (0,6) differs only a little from what is reported by Freundl and colleges (0,7).”
Abstract - supportedReviewer 2The additional use of the DaysyView app leads to higher user engagement and better understanding.The claim is based on direct self-reported survey responses and is supported as stated.Evidence: Survey responses: 84% reported better understanding, and 64.66% used the app daily.
“84% of the participants indicated that they achieved a better understanding of themselves and their cycle through the additional use of the app DaysyView.”
Results
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- N/ASurrogate endpointThe primary endpoint is unintended pregnancy, a hard clinical outcome, not a surrogate.
“2 women indicated that they had been unintentionally pregnant during the use of the device, giving a typical-use related Pearl-Index of 1.3.”
- ADEQUATEEffect sizeThe reported Pearl Indices (typical-use 1.3, method-related 0.6, perfect-use 0.8) are standard contraceptive efficacy measures and are compared to previous studies and other methods, indicating clinical meaningfulness.
“typical-use related Pearl-Index of 1.3... method-related Pearl-Index of 0.6... perfect-use Pearl-Index of 0.8.”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
7 integrity concerns flagged (0 high).
- lowmethod result mismatchThe abstract states the typical-use Pearl-Index improved 'significantly' from 3.8 to 1.3, but no statistical test comparing the two historical cohorts is reported.
“However, if the focus is on the typical-use related Pearl-Index , it has significantly improved from 3,8 to 1,3.”
Abstract
Reporting gaps
4 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data and code not sharedAssessed
- Statistical reporting inadequateAssessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
- Study-design details incomplete (controls, blinding, power)Assessed
The paper cites prior studies on fertility monitors (Freundl et al.) and apps (Setton et al.), discusses the role of apps in health behavior, and acknowledges limitations of prior app research. A logical rationale links app usage to improved usability and pregnancy rates. The hypothesis follows from the cited evidence.
“In their retrospective clinical trial, Freundl, et al., concluded that the fertility monitors Babycomp and Ladycomp achieved a method-related Pearl-Index (PI) of 0.7 and a typical-use related PI of 3.8 over 12 months”
“The aim of this study was to investigate if by the additional use of an App and thereby improved usability of the medical device, it is possible to enhance the typical-use related as well as the method-related pregnancy rates.”
“This combination is interesting because as described above, it is shown in various studies that the use of apps is increasing patients´ focus on their disease or their health behavior.”
“For example, in their retrospective clinical trial, Freundl, et al., concluded that the fertility monitors Babycomp and Ladycomp achieved a method-related Pearl-Index (PI) of 0.7 and a typical-use related PI of 3.8 over 12 months”
“However, until the present time, no study has been reported considering the contraceptive effectiveness of a fertility monitor (Daysy) optionally connected with an App (DaysyView).”
“For example, in their retrospective clinical trial, Freundl, et al., concluded that the fertility monitors Babycomp and Ladycomp achieved a method-related Pearl-Index (PI) of 0.7 and a typical-use related PI of 3.8 over 12 months”
“This combination is interesting because as described above, it is shown in various studies that the use of apps is increasing patients´ focus on their disease or their health behavior.”
Inclusion/exclusion criteria are reported (e.g., use for >13 cycles for PI calculation, exclusion of 5 participants with missing serial numbers). However, no a priori power analysis is provided, and outlier handling is not described beyond the exclusion of those 5 cases. These are common gaps in retrospective survey studies.
“In five cases, the correctness (due to the lack of the serial number of the device) of the data could not be confirmed, these participants were excluded from the study.”
“In five cases, the correctness (due to the lack of the serial number of the device) of the data could not be confirmed, these participants were excluded from the study.”
“668 respondents (2674 cycles in total) declared they had been using the fertility monitor for < 13 cycles (Fig. ). Their data was not taken into account when calculating the PI.”
The study is on women (sex is inherent), and age, BMI, cycle length, and usage patterns are reported. Demographics such as age distribution, BMI, and cycle regularity are provided. No justification for single-sex is needed as the device is female-specific.
“the fertility monitor was mainly used to avoid pregnancy (74.68% see Fig. )”
“In the resultant group of 125 women (2076 cycles in total), 2 women indicated that they had been unintentionally pregnant”
“The average age of participants was 29 years (Fig. ), whereby the fertility monitor was mainly used to avoid pregnancy (74.68% see Fig. ).”
“The average age of participants was 29 years (Fig. ), whereby the fertility monitor was mainly used to avoid pregnancy (74.68% see Fig. ).”
“The average cycle length of all participants was 28.9 (± 3.52 SD) days.”
The study was approved by a named ethics committee (FAU/Erlangen/276_16B), informed consent was obtained from all participants, and the study was conducted in accordance with the Declaration of Helsinki.
“The study protocol was reviewed and authorized by the regional ethics committee (FAU/ Erlangen/ 276_16B).”
“All patients gave informed consent to participate to the study and to publish the study data.”
“The study has been performed in accordance with the Declaration of Helsinki”
“The study has been performed in accordance with the Declaration of Helsinki and has been approved by the regional ethics committee (FAU/ Erlangen/ 276_16B).”
“All patients gave informed consent to participate to the study and to publish the study data.”
“The study has been performed in accordance with the Declaration of Helsinki”
“The study has been performed in accordance with the Declaration of Helsinki and has been approved by the regional ethics committee (FAU/ Erlangen/ 276_16B).”
“All patients gave informed consent to participate to the study and to publish the study data.”
The fertility monitor Daysy is described with manufacturer (Valley Electronics AG) and model. The app DaysyView is also named. However, no statistical software (e.g., SPSS, R) is mentioned for the analysis. Antibodies, cell lines, and other biological resources are not applicable.
“The medical device Daysy (Valley Electronics AG, Zurich, Switzerland)”
“The medical device Daysy (Valley Electronics AG, Zurich, Switzerland) is an electronic device”
“DaysyView is a free mobile app that augments the Daysy fertility monitor.”
“The medical device Daysy (Valley Electronics AG, Zurich, Switzerland) is an electronic device that also exploits the described relationship between the menstrual cycle and fluctuations in body temperature”
Table 2 reports confidence intervals where the lower limit exceeds the upper limit for all cycles, which is a clear arithmetic error. Additionally, the paper does not verify test assumptions (e.g., normality for t-tests) and does not identify the statistical software used. These issues are compounded by the absence of individual data points and inadequate description of error bars.
“****t-test p = 0.0001”
“| 1 | 696 | 4 | 0.57 | 1.57 | 0.18 |”
“From 798 participating women, 524 (64%) indicated an additional contraceptive use”
“In the present study, the second method, based on cycles, was used to calculate the PI due to the fact that the participants supplied information on the number of cycles.”
“The same value increases significantly to 10.82% probability if a woman is considered to have had unprotected intercourse on red (fertile) as well as on green (infertile) days (Fig. imperfect use).”
The only applicable criterion is the data availability statement, which is reported but inadequate (bare 'on request' without conditions). No repository deposit, accession numbers, or code sharing are provided, as these are not applicable to the study type.
“Please contact author for data request.”
“Please contact author for data request.”
“Please contact author for data request.”
The methods section describes the device, study design, and statistical analysis in sufficient detail. All pre-specified outcomes (pregnancy rates, app usage) are reported. Limitations are discussed in a dedicated section. Conclusions are appropriately cautious. Funding sources and conflicts of interest are disclosed. However, no reporting guideline (e.g., STROBE) is referenced.
“Limitations of the study The retrospective design is a very time efficient and elegant way of answering new questions with existing data.”
“This study was funded by the Valley Electronics AG, Zurich, Switzerland. NvdR is an internal scientist and employee of the company.”
“The primary disadvantage of the retrospective study design is the limited control the researchers have over the data collection.”
“This study was funded by the Valley Electronics AG, Zurich, Switzerland.”
“The retrospective design is a very time efficient and elegant way of answering new questions with existing data. The primary disadvantage of the retrospective study design is the limited control the researchers have over the data collection.”
“This study was funded by the Valley Electronics AG, Zurich, Switzerland. NvdR is an internal scientist and employee of the company.”
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 29 references by DOI: 24 verified — 5 no DOI (shown, not verified).
- NO DOICycle monitors and devices in natural family planningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICalculation of the Pearl Index of Lady-Comp, Baby-Comp and Pearly cycle computers used as a contraceptive methodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe performance of fertility awareness-based method apps marketed to avoid pregnancyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFactors in human fertility and their statistical evaluationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIContraceptive failure of the ovulation method of periodic abstinenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
1 finding · worst lowWording, consistency and formatting errors that need correcting before submission.
- Wording or formatting errors that need correctingAssessed
22 copyedit issues flagged (2 major): mostly consistency, typo, clarity.
- MAJORconsistencyTable 2“CI, lower Limit (%) | CI, upper Limit (%) | ... 1 | 696 | 4 | 0.57 | 1.57 | 0.18”→ The lower limit should be less than the upper limit. For cycle 1, likely 0.18 to 1.57.Confidence intervals are reversed; this is a data presentation error.
- MAJORconsistencyTable 2“CI, lower Limit (%) | CI, upper Limit (%)”→ Swap the two columns so that the lower limit is smaller than the upper limit for every row.All rows list a larger value under 'lower' than under 'upper' (e.g., 1.57 vs 0.18), which is impossible for a confidence interval.
- MINORtypoAbstract, line 2“trough the additional use of a mobile application”→ through the additional use of a mobile applicationTypo: 'trough' should be 'through'.
- MINORtypoPlain English summary“trough the additional use of a mobile application”→ through the additional use of a mobile applicationSame typo as above.
- MINORtypoResults, paragraph 1“the total number of recorded cycles was 4738”→ the total number of recorded cycles was 4750 or clarify the discrepancyLater in the Results, 4750 cycles are mentioned. This inconsistency needs resolution.
- MINORgrammarDiscussion, paragraph 3“the fertility monitor is short time on the market”→ the fertility monitor has been on the market for a short timeAwkward phrasing.
- MINORclarityDiscussion, paragraph 4“inperfect-use”→ imperfect-useTypo in figure label.
- MINORconsistencyMethods, Aim of the study“feasability”→ feasibilityTypo.
- MINORotherAbbreviations“n/c Not significant”→ n/s Not significant (consistent with text) or define n/sThe text uses 'n/s' but the abbreviation list says 'n/c'. Inconsistency.
- MINORclarityMethods, Statistical analysis“the app offers the opinion to share cycle information”→ the app offers the option to share cycle informationTypo: 'opinion' should be 'option'.
- MINORtypoPlain English summary“trough the additional use”→ through the additional useSpelling error.
- MINORtypoPlain English summary“Freundl and colleges”→ Freundl and colleaguesMisspelling.
- MINORtypoMethods, The app DaysyView“The app offers the opinion to share cycle information”→ The app offers the option to share cycle informationWord choice error.
- MINORconsistencyResults, paragraph 2“524 (64%) indicated an additional contraceptive use”→ 524/798 = 65.7%; correct the percentage.Percentage does not match the reported counts.
- MINORconsistencyResults, app usage“516 out of 798 (64.66%) ... 239 (29.24%) ... 44 (5.51%)”→ Check counts and percentages; 516+239+44 = 799 and 239/798 = 29.95%.Counts and percentages are internally inconsistent.
- MINORconsistencyResults, account status“776 (98%) ... 20 (2%) ... Of the 778 remaining accounts”→ Reconcile 776 + 20 = 796 with 798 participants and 778 remaining accounts.Numbers do not add up.
- MINORclarityAbbreviations“n/c Not significant”→ Use 'n/s' consistently, as the text elsewhere uses 'n/s'.Inconsistent abbreviation.
- MINORtypoAbstract, Plain English summary“trough the additional use of a mobile application”→ through the additional use of a mobile applicationTypo: 'trough' should be 'through'.
- MINORgrammarAbstract, Background“It is reported, that trough the additional use of a mobile application the interest and motivation of a patient’s health behavior increases significantly.”→ It is reported that through the additional use of a mobile application, the interest and motivation of a patient's health behavior increases significantly.Comma after 'reported' is unnecessary; 'trough' typo; missing comma after 'application'.
- MINORconsistencyMethods, Statistical analysis“The perfect-use as well as the method and typical-use related pregnancy rate (PI) was calculated separately.”→ The perfect-use, method, and typical-use related pregnancy rates (PI) were calculated separately.Subject-verb agreement and clarity.
- MINORclarityResults, paragraph 3“From a total number of 798 women using the fertility monitor for family planning, contraception or both, a total of 4750 cycles were identified.”→ From a total of 798 women using the fertility monitor for family planning, contraception, or both, 4750 cycles were identified.Simplify wording.
- MINORpunctuationDiscussion, paragraph 3“If the typical-use related PI of 6.9 is considered, it becomes clear that the user in itself represents the greatest risk (which is our main hypothesis).”→ If the typical-use related PI of 6.9 is considered, it becomes clear that the user herself represents the greatest risk (which is our main hypothesis).'in itself' should be 'herself' for a female user.
This published paper has several issues that an informed reader should weigh: the reversed confidence intervals in Table 2 and internal inconsistencies in reported counts suggest errors in data presentation that warrant a correction or erratum. The lack of a specific data availability statement and statistical software identification reduces reproducibility. The claim of 'significant' improvement relative to historical cohorts is not supported by a formal statistical comparison. A reader should interpret the reported pregnancy rates with caution and consider an independent re-analysis of the raw data if obtainable.
- 1.HIGHstatisticsCorrect Table 2 by swapping the 'CI, lower Limit' and 'CI, upper Limit' columns so that the lower limit is smaller than the upper limit in every row.The current confidence intervals are impossible (lower > upper) and constitute a data presentation error that undermines the reported pregnancy probabilities.
- 2.HIGHstatisticsRecompute and correct all percentages in the Results to match the reported counts (e.g., 524/798 = 65.7% not 64%; 239/798 = 29.95% not 29.24%; 69/798 = 8.65% not 9.01%).Several percentages are inconsistent with the reported numerators and denominators, indicating calculation errors that need resolution.
- 3.HIGHstatisticsResolve the discrepancy between '4738 cycles' and '4750 cycles' reported in the Results.The total number of cycles is reported inconsistently, which affects the denominator for Pearl Index calculations.
- 4.HIGHstatisticsReconcile the inconsistent participant counts: 524 vs 493 additional-contraceptive users, and 776+20=796 vs 798 participants / 778 remaining accounts.Internal contradictions in participant flow and subgroup counts confuse the analysis population.
- 5.HIGHdata codeReplace the vague data availability statement ('Please contact author for data request') with a concrete managed-access route (e.g., a named data access committee, institutional repository, or DOI with conditions and timeframe).The current statement provides no mechanism or timeline, making the data effectively inaccessible and hindering reproducibility.
- 6.HIGHstatisticsReport the statistical software used (e.g., SPSS, R, SAS) with version number in the Methods section.Missing software identification reduces the reproducibility of the analyses.
- 7.HIGHstatisticsAdd a statement about verification of test assumptions (e.g., normality for t-tests, proportional hazards for Kaplan-Meier) in the Statistical analysis subsection.Without assumption verification, the validity of the reported inferential tests is uncertain.
- 8.HIGHstatisticsProvide exact p-values (e.g., 'p=0.04') instead of threshold statements like 'n/s' or 'p<0.05' for all hypothesis tests.Threshold-only p-values are insufficient for readers to assess the evidence strength.
- 9.HIGHreportingTemper the claim that the typical-use Pearl-Index 'significantly improved' from 3.8 to 1.3, as no statistical test comparing the two historical cohorts is reported.The word 'significantly' implies a formal comparison that was not performed; the claim should be framed as an observed difference without inferential statistics.
- 10.HIGHreportingAcknowledge that the observed improvement in pregnancy rates with the app is an association, not a causal effect, due to the lack of a control group and reliance on historical comparators.The claim that combining Daysy with DaysyView improves usability and pregnancy rates is overstated given the study design limitations.
- 11.MEDIUMreportingAdd a completed STROBE reporting checklist for this observational study in the supplementary material.Adherence to recognized reporting guidelines improves transparency and is expected by many journals.
- 12.MEDIUMcopyeditFix the typo 'trough' to 'through' in the Abstract and Plain English summary.Spelling errors reduce professionalism and readability.
- 13.MEDIUMcopyeditCorrect the abbreviation list: 'n/c Not significant' should be 'n/s' to match the text's usage.Inconsistent abbreviations confuse readers.
- 14.MEDIUMcopyeditClarify the phrase 'the fertility monitor is short time on the market' to 'the fertility monitor has been on the market for a short time'.Awkward phrasing detracts from clarity.
- 15.LOWreportingPre-specify and report the inclusion/exclusion criteria (including the <13-cycle exclusion) as an analysis plan, and describe how missing or incomplete survey data were handled.A priori specification of criteria strengthens the study design and reduces post-hoc decisions.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.